This is basically the answer, they generate A LOT of synthetic task rollouts in parallel, then use RL on the resulting reward signals to improve the model. Add scale to this and you have a Fable class model.
keep in mind fable = mythos which as been "done" since february. so the gap is not 2 months, it's more like - techniques probably started "working" in late 2025, now are trickling down to 2nd tier labs 9 months later.
Yes it does, it just means all the companies come out with similar models around the same time. If what they were doing was completely novel, it would take a long time to repeat. As it is now each company releases a new model every few months, and every couple years the "leading" company changes.
Who are these task producers? Are you saying that Anthropic, et al delegate the RL part to third party companies that do it for pretty much every other AI company as well?
Yes they're called RL gym companies and there's a whole ecosystem of them. You hardly hear about them because their only customers are AI labs and RLVR is where the improvements are coming from at the frontier right now.
Note that RLVR is incredibly compute expensive but it's CPU as much as GPU.