RL EnvironmentsLong-Horizon EvalsAI Training Gyms
We build reinforcement-learning environments. Off the shelf, or built to your spec. Labs train in them and run evals that don’t saturate the week after they ship.
Coding from real repos. Knowledge work from real companies. And the agent fundamentals under both.
A researcher cannot improve a model against a metric they cannot trust.
Most of what is sold today measures capability, and measures it loosely enough that an agent can win without doing the work. It hacks the grader. It argues the nitpick. It learns a skill that exists nowhere outside the test. Every one of those is a reward the model collects for a solution that was never real.
The optimistic version of that failure is slow progress. The other version is a model that is genuinely superb at things that are contrived, wasteful, and wrong, and a decade spent finding out.
Intelligence is what reality selects for. Model the world, adapt across time, act through uncertainty.
So the training world has to be a piece of the real one. Every Idler environment starts from a moment when an expert created value that someone actually paid for: a bug closed in a production repo, a filing that cleared, a forecast that held. That moment becomes a task.
Then we try to break it. We solve the task end to end, sometimes a hundred times over, hunting the ambiguity that punishes a correct answer, the shortcut that rewards a wrong one, and the skill that transfers nowhere. Tasks that can be fixed get fixed. Tasks that cannot get thrown out, and most of the work is knowing which is which.
What survives is hard, fair, and real. When a model fails an Idler task, the model failed.
Bugs, features, and refactors from production codebases.
Work the economy runs on, from companies with revenue, customers, and a history.
The skills every long task decomposes into: tool use, instruction following, planning, and the long tail after them.
In the tasks, the grading, and everything around them. Hills worth climbing.
Fair by construction: every task solved end to end, many times over, before it ships. When a model fails an Idler task, the model failed.
Every task held to the frontier of its day, from the first to the ten-thousandth.
Tasks the most capable models still fail, and new ones when they stop.
03Off the shelf or built to your spec. Every layer of the environment is yours to set, from the tools an agent can call to the way the work is delivered.
Get in touchIt starts with a call. Tell us what you’re measuring, training, or trying to build.
Start a conversation