Evaluating more than speed
OmegaUse-OfficeVal, released July 29, evaluates long-horizon document, spreadsheet, and presentation work through both deliverable quality and economic value. It addresses benchmarks that report success without connecting it to human time or task cost.
One hundred tasks with economic signals
The benchmark contains 100 practitioner-proposed requests adapted through a privacy-preserving process. A human took 2.32 hours on average per task. Each task includes human labor time and a task-price proxy, allowing comparisons of labor cost, LLM inference cost, and value-weighted performance.
Fine-grained rubrics were implemented as code-based verifiers. Across several frontier LLMs and a human baseline, the researchers report that all models were faster and cheaper but still below human deliverable quality.
Metrics for deployment
Time saved is incomplete if correction and review costs are ignored. Teams should measure time to first draft, human editing through approval, rerun cost, and critical errors.
Start with drafts, organization, and format conversion where outputs are easy to verify. Final deliverables should pass deterministic checks and human review rather than relying on the agent's claim of completion.
Limitations
This preprint covers 100 tasks, selected models, and specific price assumptions. Prices, models, wages, security, and training costs vary. Because the code and data are open, organizations should rerun the evaluation with their own tasks and economics.