Characterize
Rigorous evaluations of what agents can actually compute: long‑horizon tasks, search, planning, tool use. Measurement precise enough to trust at the frontier.
We are an applied AI lab. We chart the world’s hardest computational problems and measure what agents can truly do with them.
Manifesto
The world’s hardest problems are computational: alpha hidden in noisy markets, distributed systems optimized to the microsecond, discoveries buried in exponentially large spaces.
We hunt these problems down, distill them, and measure exactly what today’s agents can and cannot do with them.
Where measured capability ends, our real work begins: moving agents deeper into problems once thought unreachable.
We define the frontier. Then we move it.
The Third Moment
μ = E[X]
The center: how capable an agent is on the typical task. Every capability story starts here, and none should end here.
σ² = E[(X−μ)²]
The spread: how consistent an agent is across runs, tasks, and days. Reliability turns raw ability into something you can build on.
μ₃ = E[(X−μ)³]
The asymmetry: how far the right tail reaches past the typical. Rare wins on brutally hard problems reveal where the frontier truly is. This is where we work.
What we do
Rigorous evaluations of what agents can actually compute: long‑horizon tasks, search, planning, tool use. Measurement precise enough to trust at the frontier.
Where agents fail, why they fail there, and how capability scales. The structure beneath the benchmark number.
Methods that move the frontier: training, scaffolding, and inference‑time compute that turn characterization into capability.
Principles
Progress moves at the pace of expert attention: hours of focus applied to problems that need years of it. The hardest problems outlast the people working on them.
Agents that search, plan, and iterate without pause turn throughput from a fixed ceiling into a quantity that scales with compute.
You cannot improve what you cannot resolve. Our evaluations are precision instruments, built to be trusted at the edge.
Characterization is the first half of the loop. Training, scaffolding, and test‑time methods close it.
Research
We are sharing them selectively while publications are in
preparation.
Contact us
to read about what we are finding.