RL EnvironmentsLong-Horizon EvalsAI Training Gyms

Push the frontier
of your model

We build reinforcement-learning environments. Off the shelf, or built to your spec. Labs train in them and run evals that don’t saturate the week after they ship.

Coding from real repos. Knowledge work from real companies. And the agent fundamentals under both.

What we build
RL Environments Gyms Evals Benchmarks
01The problem

A researcher cannot improve a model against a metric they cannot trust.

Most of what is sold today measures capability, and measures it loosely enough that an agent can win without doing the work. It hacks the grader. It argues the nitpick. It learns a skill that exists nowhere outside the test. Every one of those is a reward the model collects for a solution that was never real.

The optimistic version of that failure is slow progress. The other version is a model that is genuinely superb at things that are contrived, wasteful, and wrong, and a decade spent finding out.

02The thesis

Intelligence is what reality selects for. Model the world, adapt across time, act through uncertainty.

So the training world has to be a piece of the real one. Every Idler environment starts from a moment when an expert created value that someone actually paid for: a bug closed in a production repo, a filing that cleared, a forecast that held. That moment becomes a task.

Then we try to break it. We solve the task end to end, sometimes a hundred times over, hunting the ambiguity that punishes a correct answer, the shortcut that rewards a wrong one, and the skill that transfers nowhere. Tasks that can be fixed get fixed. Tasks that cannot get thrown out, and most of the work is knowing which is which.

What survives is hard, fair, and real. When a model fails an Idler task, the model failed.

03Frontiers

Coding

Bugs, features, and refactors from production codebases.

GDPVal

Work the economy runs on, from companies with revenue, customers, and a history.

Agentic primitives

The skills every long task decomposes into: tool use, instruction following, planning, and the long tail after them.

04Why Idler

Taste

In the tasks, the grading, and everything around them. Hills worth climbing.

Rigor

Fair by construction: every task solved end to end, many times over, before it ships. When a model fails an Idler task, the model failed.

Consistency

Every task held to the frontier of its day, from the first to the ten-thousandth.

Frontier

Tasks the most capable models still fail, and new ones when they stop.

05Off the shelf, adapted, custom
Off the shelf
Task suites, ready to run.
Adapted
Existing environments reconfigured to your requirements.
Custom
Built to your spec.

Tune every parameter

03Off the shelf or built to your spec. Every layer of the environment is yours to set, from the tools an agent can call to the way the work is delivered.

Get in touch
ToolsEnterprise backends, filesystems, custom toolchains.
DifficultyCalibrated to your target model, tuned by pass@n.
CapabilitiesDomain skills, enterprise tools, agent behaviors.
Use-caseEvaluation, training, benchmarking, in combination.
DeliveryHosted, exported, or integrated into existing pipelines.

06Working with Idler
It starts with a call. Tell us what you’re measuring, training, or trying to build.
FitIf a suite already fits, it ships.
SampleIf not, we sample tasks from your past work right away.
PilotA two-week pilot: a small set of tasks, iterated together until they hold. Flat fee.
ScaleThen we scale, from tens to tens of thousands of tasks.

Get in touch

It starts with a call. Tell us what you’re measuring, training, or trying to build.

Start a conversation