he did say that GPT-4 would use a different loss function, what effect would that have? I have no idea.
One possibility is shifting the power law. See UL2 which combines the various denoising losses in what turns out to be a very good way: “U-PaLM: Transcending Scaling Laws with 0.1% Extra Compute”, Tay et al 2022 - halving PaLM training requirements w/UL2 losses. I don’t know if OA discovered UL2 first, but it’s not all that exotic or subtle and is certainly something that many people ask themselves when they learn about the difference between bidirectional and unidirectional models: “why not train on both/all the losses?”
One possibility is shifting the power law. See UL2 which combines the various denoising losses in what turns out to be a very good way: “U-PaLM: Transcending Scaling Laws with 0.1% Extra Compute”, Tay et al 2022 - halving PaLM training requirements w/UL2 losses. I don’t know if OA discovered UL2 first, but it’s not all that exotic or subtle and is certainly something that many people ask themselves when they learn about the difference between bidirectional and unidirectional models: “why not train on both/all the losses?”