Physix Frontier · News Briefing Card (enbrief · Oct 2, 2026)
AI Agent Training Data Bottleneck: Open-Ended Task Data Missing
KEY FACTS
- The article points out that current AI agent training data is mostly structured, single-objective task data, lacking the open-ended data found in real work scenarios.
- Opus 5.5 and GPT-6 Astra claim scores of 72.6% and 81.8% respectively on OSWorld 2.0.
- Computer-use agents rely on trajectory data for training, where trajectories record the detailed steps of humans or agents operating a computer.
- Data quality scales along two dimensions, observability and distribution, with observability advancing faster than distribution.
- Real work involves interleaved multitasking, interruption recovery, and context switching, and existing data does not capture these signals.
KEY DATA
72.6%Opus 5.5 on OSWorld 2.0
81.8%GPT-6 Astra on OSWorld
PHYSIX OBSERVATION
Models post eye-catching scores on structured benchmarks, but real knowledge work is full of interruptions, multithreading, and ambiguous goals, and existing trajectory data barely covers any of it. This means agents will still stumble frequently when deployed in office scenarios. Whoever first solves the collection and annotation of open-ended work data will secure the ticket into the next stage of enterprise-grade agents.
Source: enbrief original report ↗
Physix Frontier