Let's bet on this shit
How to act rationally if you believe in RL scaling and envs

People talk a lot about acting rationally once you consider your beliefs as axiomatic, particularly in AI. We have many examples of this. For instance, people have been talking about scaling laws since GPT-2 (and even beforehand); they would tell you to be scaling pilled. Well, we didn’t actually see that many of them throwing away everything to get more compute, more data, and investing in NVIDIA call options. Similarly, people talk about being AGI pilled. Well, if that’s your belief, that AGI will hit in the next year or two with the current transformer paradigm, why aren’t you buying as much land as possible, or folding your little neolab startup and begging one of the AGI-leading labs to let you in?
I think we are probably going through a similar phase right now with data, particularly RL data. People are claiming that everything is going to be RL environments, and yet I see very few people acting rationally on this belief to its logical conclusion. Suppose you believe that the current paradigm we have will work, and that RL is its own new scaling law, and the only thing between us and RSI is better, more targeted, more RL environments. What would you rationally do?
As Adam Sandler would say, what you really should be saying is “Let’s fucking bet on this shit”.
As we get closer to RSI, the amount of compute, time, manual labour and care required to produce an environment of unit utility to the model’s capabilities is likely to increase exponentially. Because of this, useful environments get scarcer, and since RSI is essentially infinite positive EV, the value of one that actually works should increase exponentially too. (The human cost per environment going up is also exactly why you’d have models making them; more on that below.) So I’d learn to make the highest value, most RSI-targeted environments I could. And maybe you need to be inside a lab to do this, because RSI is likely cumulative and builds on the training stacks the labs have internally. I’m going to write more about that later.
(And again, I’m saying nothing about my own personal beliefs on this. I’m just pointing how the inconsistencies in other peoples’.)
And let’s suppose you didn’t even believe in RSI. I actually think this is the case where you should probably care about RL envs even more! If we don’t have RSI, then the labs are going to need to make their models as good as possible at everything, which includes niche finance tasks and powerpoints and genetic analysis and agricultural plant disease recognition and so on. Which means that they will have a very high willingness to pay for data! So make useful data. Figure out what it means to make data which makes the models more useful.
And this seems simple, I haven’t said anything non-trivial yet. Maybe what’s non-trivial is that even people working on AI research themselves haven’t really internalised this. Like, if you believe truly in the value of data, instead of doing mech interp, you’d probably just be building the environments to do mech interp. Similarly, on a somewhat meta point, if you believe in the value of RL environments, you should probably not be making RL environments, but indeed making RL environments to make RL environments. (I guess you would do this if you believed in RSI, and not otherwise.)
Another perhaps non-obvious point to make is about how valuable different types of tasks are based on the generalisability of RL. I believe the capabilities learned (well, unearthed/uplifted) during RL don’t generalise very much between domains, or even intra-domain. What I do believe generalises is the horizon length of how long a model can make meaningful progress on a given problem. So I would focus on making longer and longer tasks (where the length is coming from some irreducible amount of meaningful complexity you’re adding to the problem, not from adding arbitrary layers of data or difficulty that are ersatz difficulty) in the domains you have the most expertise in. The rest is baking in specific intuitions and taste and knowledge which got spread too thin to be useful in pretraining.
Even people starting their own RL environments companies are largely not following this. Most of them (and I’ve by now been inside enough of them to have lost my appetite for the tour) are not companies so much as Goodhart machines. They are apparatus (apparati?) for converting conviction into tonnage, a great slurry of tasks patched together overnight and honed to a strange thin edge against some synthetic failure mode that nobody had ever named before, that most humans would also cut themselves on, and that no one (including whoever built the task) could tell you how to handle a priori. If you believed everything was going to be RL environments, you would build them the way you’d build a cathedral. You wouldn’t be trying to get rich quick. You’d try and truly, fundamentally understand what makes a good RL environment, not just what the minimum criteria are for a lab to buy something off you.
A final point about how I’d specifically handle RL environments (this is the long one, so bear with me). My claim is that, at least in a world without RSI, a broad swathe of RL environments across many domains, task types and difficulty levels is worth more than a small set of very high quality, very hard environments at the frontier. (Or a broad swathe of data, at least. RL environments give you mid-training data for free, so the argument is the same either way.)
The reasoning is as follows. Anthropic can distill Sonnet 5 and Opus 5 from Mythos, and they can do it on logits. Kimi K3 and GLM 5.3 can at best do token distillation off the frontier models’ outputs. And yet Kimi K3 and GLM 5.3 are the better models. The naive read is that distillation just isn’t worth much. An alternative reading is that distillation is worth a lot, but what you distill from matters more than how well you can distill. To make this case, look at what each side is distilling from. The Chinese labs have access to router farms, which means they see real user queries and the frontier models’ real responses across a representative range of use cases. That is exactly the data token distillation needs, and it’s the real distribution of what people actually ask. Anthropic and co, as far as anyone can tell, are keeping their word about not training on user traces directly. They can summarise and inspect at a high level, but not train on them (which is partly why they ask on Twitter about their models’ failure modes). What Anthropic does have is much better environments at the frontier, because they pay a lot more for them. So one side is doing cheap distillation on the real distribution, and the other is doing expensive distillation on synthetic, hard, frontier environments. The cheap side is winning. So according to this argument an RL environment seeded from real-world user data is worth more than a weirdly synthetic, very hard one you paid Mercor or Surge or Handshake to make.
So I guess whichever branch you take, then you reason the following way. If RSI is real, the value of a unit-utility environment goes exponential and you should be making the most targeted, most cumulative ones you can, probably inside a lab. If RSI isn’t real, the labs need to be good at everything, and their willingness to pay for useful data is enormous. Either way, the environments worth building are long (because horizon is the only thing that reliably generalises), real (because what you distill from matters more than how much you can distill), and yours (because the taste and domain knowledge that got spread too thin in pretraining is exactly what’s left to put in). Not really many people are doing this properly, which is the same gap as the scaling-pilled person who never bought the NVIDIA calls and the AGI-pilled person who never bought the land.
LONG REAL YOURS