The Real-World Agent Benchmark: 100 Tasks Humans Actually Pay For
AI agents are becoming impressive at benchmarks. But businesses do not pay agents for benchmark points. They pay for completed research, reconciled accounts, qualified leads, fixed software, resolved support tickets and other useful outcomes. The DN Real-World Agent Benchmark proposes a harder test: can an agent reliably finish work someone would actually pay a human to do?
Framework: DN-RWAB 1.0 · Last verified: 29 September 2026 · Dataset design: 100 economically useful tasks across 10 work categories
What Matters
The next useful AI-agent leaderboard should not ask only whether an agent can solve a coding issue, browse a sandbox or call the correct API. It should ask whether the agent can deliver paid-quality work. DN-RWAB scores 100 practical tasks on completion, correctness, human intervention, time, cost, reliability and economic value.
DN Evidence Block
Why this exists: Current agent benchmarks provide valuable evidence about coding, browser use, computer control, tool use and task horizon. But these tests do not by themselves establish whether an agent can repeatedly perform the heterogeneous work businesses and individuals buy in the open economy.
The Signal: Benchmark Scores Are Not the Same as Economic Utility
AI benchmarking has progressed rapidly.
SWE-bench evaluates whether systems can solve genuine software-engineering issues. Its Verified dataset contains 500 tasks reviewed by humans for clarity and solvability.
OSWorld moves evaluation into real computer environments rather than static prompts. Its newer generations increasingly emphasize longer, more realistic application workflows.
Princeton's τ-bench family evaluates agents that must converse with users, invoke tools and follow policies inside realistic domains such as airline support.
METR measures task-completion time horizons by comparing agent success against the amount of time expert humans require to complete tasks.
These are meaningful advances.
But each benchmark measures a slice of useful work.
The Missing Unit: Paid-Quality Completion
Work has several characteristics that ordinary benchmark success rates can hide.
A research report can contain the requested facts and still be unusable because its sources are wrong.
A sales prospect list can contain 100 rows and still be worthless because half the contacts are irrelevant.
A spreadsheet can calculate correctly while silently corrupting the formatting or formulas elsewhere.
A customer-service agent can eventually reach the correct outcome but consume more time and money than a trained employee.
An agent can successfully finish a task once and fail the next four times.
That means binary completion is insufficient.
DN proposes a higher standard:
The DN 100-Task Real-World Agent Benchmark
DN-RWAB 1.0 divides 100 tasks across ten categories representing commercially meaningful knowledge work.
| Category | Tasks | Example paid work | Primary failure mode |
|---|---|---|---|
| Research & Analysis | 10 | Competitive research, source verification, industry briefings | Unsupported conclusions or missed evidence |
| Software & Technical Work | 10 | Bug fixes, scripts, debugging, data transformations | Correct-looking output that breaks elsewhere |
| Finance & Accounting | 10 | Reconciliation, categorization, variance analysis, invoice checks | Silent numerical or classification errors |
| Sales & Lead Generation | 10 | Prospect research, qualification, outreach preparation | Volume without relevance or accuracy |
| Marketing & Content | 10 | Campaign research, briefs, SEO analysis, publishing workflows | Plausible but generic output |
| Customer Operations | 10 | Support resolution, refunds, account changes, escalation | Policy violations or incorrect actions |
| Commerce & Procurement | 10 | Product sourcing, quote comparison, purchasing decisions | Wrong specifications or hidden total cost |
| Administrative Work | 10 | Scheduling, document processing, inbox triage, travel planning | Missed constraints or incomplete execution |
| Data & Spreadsheet Work | 10 | Cleaning, analysis, formulas, reports, dashboard preparation | Subtle data corruption |
| Agentic Finance | 10 | Route comparison, treasury checks, transaction preparation, risk monitoring | Financial loss or policy breach |
What the 100 Tasks Should Actually Look Like
A useful benchmark should avoid toy requests that can be solved by pattern matching.
Each DN task should contain realistic constraints, imperfect information and a verifiable outcome.
Research
Identify five operational competitors to a specified company, verify their current offerings from primary sources and deliver a comparison with dated evidence.
Spreadsheet
Clean a messy transaction export, preserve source data, categorize expenses and produce a reconciled summary without breaking formulas.
Customer Support
Resolve a customer's account problem while following refund, privacy and escalation policies.
Procurement
Compare suppliers for an exact specification while incorporating delivery, minimum order quantities and total landed cost.
Software
Diagnose a real failure, implement a fix, verify regressions and explain the change in language another developer could review.
Finance
Reconcile payments against invoices and flag exceptions that require human investigation rather than inventing explanations.
DN's Seven-Dimensional Scoring Model
Every benchmark run should generate more than a pass or fail.
| Metric | Weight | What it measures |
|---|---|---|
| Task completion | 25% | Whether the requested outcome was actually achieved |
| Correctness | 20% | Whether facts, calculations and actions were materially accurate |
| Output usability | 15% | Whether the deliverable could be used without substantial repair |
| Reliability | 15% | Whether repeated runs produce consistently successful outcomes |
| Human intervention | 10% | How much rescue, clarification or correction was required |
| Cost efficiency | 10% | Agent cost compared with the estimated human cost of completing the same task |
| Time efficiency | 5% | Elapsed completion time relative to a competent human baseline |
Try the DN Economic Task Score
DN Economic Task Calculator
Estimate whether an agent output qualifies as economically useful work rather than merely technical completion.
Not Yet Economically Useful
Select the observed result for each dimension.
Interpretation- Measure the finished outcome, not the agent's reasoning style.
- Repeat the task before drawing production conclusions.
DN Economic Task Score Bands
| Score | Classification | Meaning |
|---|---|---|
| 0–39 | Not Economically Useful | The agent may demonstrate capability but cannot yet substitute for paid work. |
| 40–59 | Human-Dependent | Useful as assistance, but human correction materially contributes to the final outcome. |
| 60–79 | Commercially Useful | The agent can complete meaningful work but still needs defined review and failure handling. |
| 80–100 | Paid-Quality Autonomous | The agent repeatedly produces usable results at commercially compelling cost and reliability. |
The Benchmark Must Penalize Lucky Runs
One successful demonstration is not reliability.
Agent systems are probabilistic, and multi-step workflows create more opportunities for small differences to compound.
DN-RWAB therefore proposes at least three independent runs per task for baseline testing.
High-value or safety-sensitive tasks should eventually use more.
Why Repeatability Changes the Leaderboard
Imagine two agents.
Agent A successfully completes 85 of 100 tasks on its best run.
Agent B completes 78.
That headline makes Agent A appear better.
Now run each task five times.
If Agent A succeeds unpredictably while Agent B produces the same correct result almost every time, the commercial conclusion changes.
Businesses often care more about predictable 78% performance than an unstable system capable of occasionally reaching 85%.
Public Benchmarks Already Show Why Harnesses Matter
Agent evaluation measures more than the underlying model.
The same model can perform differently depending on its tools, prompts, memory architecture, environment, retry strategy and agent scaffolding.
SWE-bench explicitly distinguishes model comparisons from broader agent-system submissions. Princeton has also removed compromised benchmark results when test-set leakage was identified in a τ-bench scaffold.
That is important.
DN-RWAB Should Score the System, Not the Brand
Every leaderboard entry should record:
- model and exact version;
- agent framework or harness;
- tools available;
- maximum steps;
- reasoning or compute configuration where disclosed;
- external search access;
- memory configuration;
- human interventions;
- retries;
- total tokens or compute usage where measurable;
- API and tool cost;
- wall-clock completion time;
- test date;
- task version.
Without these details, a leaderboard risks comparing fundamentally different systems under one model name.
The Cost per Successful Task Is More Useful Than Token Price
Cheap tokens do not necessarily create cheap work.
Suppose Agent A costs $0.20 per attempt but succeeds only 30% of the time. Its effective model cost per successful task is approximately $0.67 before human rescue costs.
Agent B might cost $0.50 per attempt but succeed 90% of the time, producing an effective cost near $0.56 per successful completion.
The supposedly more expensive agent is economically better.
This is why DN proposes:
Human Intervention Must Be Priced
One of the largest hidden costs in agent deployments is human rescue.
An agent may cost pennies to run but require ten minutes of professional review after every task.
At scale, the review cost can dominate inference.
DN-RWAB therefore treats human involvement as part of the economic cost rather than pretending it is free.
| Intervention level | Example | DN treatment |
|---|---|---|
| None | Agent completes and verifies task independently | 0 rescue minutes |
| Light review | Human checks final output without materially changing it | Record review time |
| Correction | Human fixes errors before output becomes usable | Deduct usability and intervention points |
| Rescue | Human must complete a failed part of the workflow | Material score penalty |
| Takeover | Human finishes the task | Agent does not receive full completion credit |
The Economic Replacement Ratio
DN proposes another metric for the benchmark:
An Economic Replacement Ratio of 1 means agent and human cost are approximately equivalent.
A ratio of 5 means the verified agent workflow costs roughly one-fifth of the competent human baseline.
But cost advantage alone is insufficient.
A cheap workflow producing unusable work has no meaningful replacement value.
Example: A $100 Human Task vs a $5 Agent Run
Consider a market-research task that a competent freelancer would charge approximately $100 to complete.
The agent workflow costs $5.
At first glance, the economic replacement ratio appears to be 20×.
But suppose the agent produces a usable answer only half the time and requires 20 minutes of human checking and correction.
The genuine cost advantage could be dramatically lower.
Benchmarking therefore has to measure the entire production loop.
Ten Research Tasks Humans Actually Pay For
| # | Task | Verified outcome |
|---|---|---|
| 1 | Competitor landscape | Current competitors, evidence and differentiated comparison |
| 2 | Supplier research | Qualified suppliers meeting exact constraints |
| 3 | Market sizing | Transparent assumptions and reproducible calculation |
| 4 | Regulatory research | Current primary-source requirements and jurisdiction boundaries |
| 5 | Company diligence | Verified facts separated from inference |
| 6 | Literature review | Relevant papers, findings and limitations |
| 7 | Product comparison | Current features, pricing and decision-relevant differences |
| 8 | News intelligence | Material developments with publication and event dates distinguished |
| 9 | Executive briefing | Concise decision-ready synthesis |
| 10 | Evidence audit | Unsupported claims identified and corrected |
Ten Finance and Accounting Tasks
| # | Task | Verified outcome |
|---|---|---|
| 11 | Bank reconciliation | Transactions reconciled with exceptions isolated |
| 12 | Invoice matching | Correct invoice-payment mapping |
| 13 | Expense categorization | Consistent categories with ambiguous items flagged |
| 14 | Budget variance analysis | Correct variances and material drivers identified |
| 15 | Cash-flow forecast | Transparent assumptions and internally consistent projection |
| 16 | Duplicate-payment check | True duplicates isolated without excessive false positives |
| 17 | Management reporting | Correct financial summary built from supplied data |
| 18 | Collections prioritization | Receivables correctly segmented by urgency |
| 19 | Purchase-order check | PO, invoice and receipt discrepancies identified |
| 20 | Financial anomaly review | Material anomalies flagged without invented explanations |
Ten Sales Tasks
| # | Task | Verified outcome |
|---|---|---|
| 21 | ICP company sourcing | Companies genuinely matching defined criteria |
| 22 | Lead qualification | Relevant prospects prioritized accurately |
| 23 | Decision-maker research | Correct role and current employment verified |
| 24 | Account research | Useful trigger events and pain points identified |
| 25 | Meeting preparation | Accurate briefing with no fabricated personal details |
| 26 | CRM cleanup | Duplicates and malformed records corrected safely |
| 27 | Opportunity prioritization | Pipeline ordered using stated evidence |
| 28 | Proposal research | Proposal tailored to verified buyer requirements |
| 29 | Follow-up drafting | Relevant follow-up grounded in actual interaction |
| 30 | Lost-deal analysis | Patterns separated from speculation |
Ten Marketing and Content Tasks
| # | Task | Verified outcome |
|---|---|---|
| 31 | SEO competitor analysis | Real content gaps supported by search evidence |
| 32 | Content brief creation | Publishable brief aligned with user intent |
| 33 | Source-backed article | Accurate claims with current sources |
| 34 | Campaign analysis | Correct performance interpretation from supplied metrics |
| 35 | Social repurposing | Channel-appropriate derivatives preserving meaning |
| 36 | Newsletter production | Accurate concise edition built from source material |
| 37 | Landing-page audit | Specific conversion issues identified rather than generic advice |
| 38 | Editorial fact check | Unsupported or stale claims surfaced |
| 39 | Keyword clustering | Search terms grouped by meaningful intent |
| 40 | Content refresh | Stale claims updated without degrading existing useful material |
Ten Software and Technical Tasks
| # | Task | Verified outcome |
|---|---|---|
| 41 | Bug diagnosis | Root cause correctly identified |
| 42 | Bug repair | Issue fixed without regression |
| 43 | Data-conversion script | Correct output across test cases |
| 44 | API integration | Working implementation with failures handled |
| 45 | Test creation | Tests catch intended failure modes |
| 46 | Documentation update | Documentation accurately matches implementation |
| 47 | Dependency migration | Migration completed without hidden breakage |
| 48 | Performance diagnosis | Bottleneck supported by evidence |
| 49 | Security configuration review | Material misconfigurations correctly identified |
| 50 | Deployment troubleshooting | Production issue resolved and verified |
Ten Customer Operations Tasks
| # | Task | Verified outcome |
|---|---|---|
| 51 | Order-status resolution | Correct current status communicated |
| 52 | Refund processing | Policy-compliant refund outcome |
| 53 | Account-access issue | Problem solved without bypassing security controls |
| 54 | Subscription change | Correct plan change executed |
| 55 | Complaint handling | Issue resolved or appropriately escalated |
| 56 | Billing dispute | Evidence checked before action |
| 57 | Technical support triage | Correct category and next action |
| 58 | Policy interpretation | Policy applied consistently |
| 59 | Case summarization | Complete accurate handoff note |
| 60 | Escalation decision | High-risk case escalated at correct threshold |
Ten Commerce and Procurement Tasks
| # | Task | Verified outcome |
|---|---|---|
| 61 | Product sourcing | Products satisfy all mandatory requirements |
| 62 | Quote normalization | Supplier quotes converted to comparable terms |
| 63 | Total-cost comparison | Freight, taxes, minimums and extras included |
| 64 | Vendor verification | Supplier existence and claims checked |
| 65 | Purchase recommendation | Recommendation follows stated constraints |
| 66 | Inventory exception | Shortage correctly identified and routed |
| 67 | Contract comparison | Material commercial differences extracted accurately |
| 68 | Renewal review | Pricing and contractual changes surfaced |
| 69 | Alternative supplier search | Viable substitutes meeting specification |
| 70 | Purchase-order preparation | Correct quantities, pricing and supplier details |
Ten Administrative Tasks
| # | Task | Verified outcome |
|---|---|---|
| 71 | Calendar coordination | Meeting scheduled without conflict |
| 72 | Travel planning | Itinerary respects time, price and travel constraints |
| 73 | Inbox triage | Important messages prioritized correctly |
| 74 | Document organization | Files classified without loss or duplication |
| 75 | Meeting-note extraction | Decisions and actions captured correctly |
| 76 | Form completion | Fields correctly populated from evidence |
| 77 | Deadline tracking | Dates and dependencies correctly identified |
| 78 | Expense submission | Receipts matched and policy rules respected |
| 79 | Contact-data cleanup | Records normalized without merging distinct people |
| 80 | Action follow-up | Outstanding commitments accurately surfaced |
Ten Data and Spreadsheet Tasks
| # | Task | Verified outcome |
|---|---|---|
| 81 | Dataset cleaning | Errors corrected without deleting valid variation |
| 82 | Formula repair | Broken formulas fixed consistently |
| 83 | Pivot analysis | Correct aggregation and grouping |
| 84 | Dashboard preparation | Metrics correspond exactly to source data |
| 85 | Duplicate detection | True duplicates isolated with low false positives |
| 86 | Outlier review | Potential anomalies identified without automatic deletion |
| 87 | Data merge | Datasets joined without record corruption |
| 88 | CSV transformation | Target schema produced accurately |
| 89 | Metric calculation | Formula and final value both verifiable |
| 90 | Trend analysis | Conclusions supported by underlying data |
Ten Agentic Finance Tasks
| # | Task | Verified outcome |
|---|---|---|
| 91 | Execution-route comparison | Best route identified using real total execution cost |
| 92 | Stablecoin transfer planning | Network, fees and destination compatibility verified |
| 93 | Treasury-balance review | Balances and exposures correctly summarized |
| 94 | Wallet-permission audit | Material permissions and risks identified |
| 95 | Transaction simulation | Expected state change correctly represented |
| 96 | Yield comparison | Net yield separated from headline APY and risk |
| 97 | Protocol-risk review | Current material risks supported by evidence |
| 98 | Position monitoring | Threshold breach correctly detected |
| 99 | Reconciliation | On-chain activity reconciled to internal records |
| 100 | Controlled transaction preparation | Correct transaction prepared within explicit policy limits |
The Hardest Part Is Ground Truth
Real-world tasks are more difficult to benchmark than multiple-choice questions because the correct result may contain judgment.
DN-RWAB should therefore use three types of evaluation.
Deterministic
Use exact checks wherever possible: balances, formulas, successful transactions, code tests, record states and database outputs.
Rubric Based
Use explicit criteria for outputs such as research briefs where several valid answers may exist.
Expert Reviewed
Use blinded human reviewers for tasks requiring professional judgment, with disagreements documented.
The Human Baseline Matters
An agent cannot be described as economically competitive without a comparison point.
Every task should therefore estimate:
- competent human completion time;
- competent human cost;
- expected human error range where measurable;
- minimum acceptable professional quality.
This does not mean humans must achieve 100%.
In fact, comparing agents against realistic human performance rather than an imaginary perfect worker makes the benchmark more useful.
Task Horizon Is Becoming a Critical Metric
METR's task-horizon work provides an important insight: agent capability is not only about whether a system can do something, but how long and complex a human-equivalent task it can complete reliably.
Real employment contains tasks ranging from two-minute lookups to projects that consume days.
DN-RWAB should therefore record estimated competent-human completion time for every task.
| Human task duration | DN class | Examples |
|---|---|---|
| Under 15 minutes | Micro | Lookup, categorization, simple update |
| 15–60 minutes | Short | Research note, reconciliation, lead qualification |
| 1–4 hours | Professional | Detailed analysis, software repair, data workflow |
| 4–8 hours | Workday | Major report, system migration, complex diligence |
| 8+ hours | Long Horizon | Multi-stage professional project |
Why OSWorld 2.0 Matters to This Thesis
The direction of recent benchmark design already points toward longer real-world workflows.
OSWorld 2.0 introduced 108 long-horizon computer-use workflows whose human completion time had a median of roughly 1.6 hours.
That is materially closer to professional work than a short GUI task.
Its published research still found frontier systems struggling with the full workflows, especially when success required tracking hidden state, changing information and multiple constraints.
The conclusion is not that agents are incapable.
It is that realistic work exposes failure modes that simpler tasks can miss.
The Benchmark Should Include Failure Recovery
Humans encounter broken links, missing data, failed APIs, ambiguous requests and changed circumstances.
Real agents will too.
A practical benchmark should therefore deliberately include recoverable failures.
Examples:
- a supplier page becomes unavailable;
- a spreadsheet contains an unexpected column;
- a user provides conflicting instructions;
- a tool call fails temporarily;
- a booking option disappears;
- a transaction simulation fails;
- a source contradicts another source;
- required information is missing.
A robust agent should know whether to retry, use another route, ask for clarification, escalate or stop.
The Benchmark Should Penalize Confident Guessing
One particularly expensive agent failure is plausible fabrication.
An agent that asks for missing information may seem less autonomous than one that invents an answer and proceeds.
Economically, the cautious agent may be far more useful.
DN-RWAB should therefore reward appropriate abstention.
If success requires information the system does not possess, the correct behavior may be to stop and request it.
Agents Should Be Measured on Net Labor Saved
Gross automation can be misleading.
A workflow may eliminate 60 minutes of manual work while creating 25 minutes of checking, correction and system maintenance.
The real productivity gain is smaller.
This makes it possible to compare systems that automate different portions of the same workflow.
The DN Agent Economic Frontier
Over time, the most revealing visualization may not be a single ranking.
It may be a frontier.
Plot each tested system by:
- paid-quality completion rate;
- cost per successful task;
- human intervention minutes;
- median task horizon;
- repeatability;
- failure severity.
Different systems may dominate different parts of this frontier.
One agent might excel at short low-cost administrative work.
Another may be expensive but capable of completing long technical projects.
A third may be exceptionally reliable in financial workflows where correctness matters more than raw speed.
Why a Single “Best AI Agent” Is Usually the Wrong Question
Agent performance is task dependent.
A coding agent should not automatically be considered superior to a customer-service agent because it performs better on a software benchmark.
DN-RWAB should therefore publish:
- overall score;
- category scores;
- task-horizon distribution;
- cost per successful task;
- reliability;
- human-intervention requirement.
Readers can then select systems according to the work they actually want performed.
Proposed DN Leaderboard Structure
| Agent | Paid-Quality Completion | Reliability | Median Cost / Success | Human Minutes | Longest Reliable Task |
|---|---|---|---|---|---|
| Agent A | Awaiting benchmark run | — | — | — | — |
| Agent B | Awaiting benchmark run | — | — | — | — |
| Agent C | Awaiting benchmark run | — | — | — | — |
DN will not fabricate benchmark scores before controlled runs are completed. Initial leaderboard cells should remain explicitly unscored until reproducible test data exists.
How DN-RWAB Differs From Existing Agent Benchmarks
| Benchmark | Primary question | DN-RWAB adds |
|---|---|---|
| SWE-bench | Can the system fix real software issues? | Cross-industry paid work, cost and human-intervention accounting |
| OSWorld | Can the agent operate real computer interfaces? | Economic usefulness across professions and channels |
| τ-bench | Can an agent interact with users and tools while following domain rules? | Broader work categories and explicit labor economics |
| METR Time Horizon | How long a human-equivalent task can an agent complete at a given reliability? | Task value, output usability and cost per successful outcome |
| DN-RWAB | Can this system reliably produce work someone would pay for? | Economic completion as the central unit |
For Individuals: Measure What You Can Stop Paying Someone Else to Do
For freelancers, creators and professionals, the useful question is not how many benchmarks a model tops.
It is which tasks can now be delegated without increasing mistakes.
Start by documenting recurring paid or time-consuming work:
- research;
- data cleaning;
- report preparation;
- prospecting;
- bookkeeping preparation;
- content repurposing;
- administration.
Test the agent repeatedly and compare the net labor saved.
For Businesses: Build a Private Version of the Benchmark
Public benchmarks should be treated as evidence, not production guarantees.
A company deploying agents should construct a private evaluation set from its own recurring work.
Twenty authentic internal tasks may reveal more about deployment readiness than hundreds of unrelated public benchmark tasks.
The strongest approach is therefore:
- use public benchmarks to shortlist systems;
- use DN-RWAB-style economic metrics to compare commercial usefulness;
- use private internal tasks before granting production authority;
- continue measuring live outcomes after deployment.
For Agent Developers: Optimize for Completed Economic Work
DN-RWAB creates a different engineering target.
Instead of optimizing only for tokens, benchmark percentage or tool-call accuracy, developers would optimize for:
- correct outcome per dollar;
- completed work per hour;
- low human rescue;
- high repeatability;
- safe abstention;
- reliable recovery;
- professional output quality.
Those metrics map more directly to adoption.
The Economic Threshold for Agent Adoption
An agent becomes commercially interesting when four conditions intersect:
1. Capability
The agent can genuinely perform the target workflow.
2. Reliability
It performs well repeatedly rather than occasionally.
3. Economics
Total cost is competitive after oversight and corrections.
4. Risk
The consequences of failure are acceptable and controlled.
Improving model intelligence addresses only the first condition.
The commercial agent economy depends on all four.
DN Methodology
Framework: DN Real-World Agent Benchmark 1.0
Objective: Measure whether autonomous or semi-autonomous AI systems can repeatedly produce economically useful outcomes comparable with work humans are paid to perform.
Proposed dataset: 100 tasks across ten work categories.
Primary metrics:
- 25% task completion;
- 20% correctness;
- 15% output usability;
- 15% reliability;
- 10% human intervention;
- 10% cost efficiency;
- 5% time efficiency.
Repeat testing: Minimum three independent runs per system-task combination in the first dataset release, with additional repetitions preferred for high-variance tasks.
Human baselines: Each task should include estimated competent-human completion time and cost, based on documented professional rates or commissioned human runs where practical.
System identification: Scores apply to the tested model, scaffold, tools, configuration and environment together.
Evidence boundary: The 100-task framework in this article is a benchmark specification. DN is not presenting invented performance results. Leaderboard rankings should be published only after controlled runs are completed.
Update cadence: Dataset versions should be permanent and reproducible. New task releases should use version numbers rather than silently changing historical tests.
Commercial independence: Sponsorship or affiliate availability must not change task selection, benchmark scoring or published results.
Falsification test: DN-RWAB should be revised if its score fails to correlate with independently measured real-world usefulness, economic savings or successful production deployment.
Limitations
No 100-task benchmark can represent the full labor market.
Work quality can be subjective, human cost varies by geography, and some jobs depend on tacit knowledge, relationships or physical actions that digital agents cannot reproduce.
Benchmark systems can also improve specifically against known datasets.
DN should therefore rotate private holdout tasks, publish methodology transparently and avoid presenting one aggregate score as universal intelligence.
Primary Sources and Benchmark References
Human-validated real-world software-engineering tasks used to evaluate coding agents.
SWE-bench Verified
Benchmarking multimodal agents inside real computer environments, including the newer long-horizon OSWorld 2.0 direction.
OSWorld
Evaluation of realistic tool-agent-user interactions and agent reliability.
τ-bench Airline
Measurement of the human-equivalent task duration frontier AI agents can complete at specified reliability levels.
METR Time Horizons
Evidence on the growing use of agents for longer-horizon workplace tasks and delegated work.
How Agents Are Transforming Work
Find the Right Agentic Stack
Use DN Pathfinder to compare the platforms and infrastructure relevant to what you actually want an agent to accomplish.
Open DN PathfinderFrequently Asked Questions
What is the DN Real-World Agent Benchmark?
The DN Real-World Agent Benchmark is a proposed 100-task evaluation framework designed to measure whether AI agents can repeatedly complete economically useful work that humans are currently paid to perform.
How is DN-RWAB different from other AI agent benchmarks?
Existing benchmarks often focus on a particular capability such as coding, browser use, tool calling or computer control. DN-RWAB adds cross-industry tasks and explicitly measures cost, human intervention, repeatability and paid-quality output.
What does paid-quality completion mean?
Paid-quality completion means that an agent has not merely technically finished a task. Its output must be materially correct, usable without substantial repair and sufficiently reliable that someone could plausibly pay for the result.
Why does the benchmark use repeated runs?
AI-agent behavior can vary between runs. Repetition helps distinguish a reliable system from one that occasionally succeeds by chance.
What is cost per successful task?
Cost per successful task divides total agent, tool and human-intervention costs by the number of verified successful outcomes. It is often more commercially meaningful than token price alone.
Why include human intervention in agent costs?
Human review, corrections and rescue work consume labor. Ignoring this can make an apparently cheap autonomous workflow look substantially more economical than it really is.
Does a high benchmark score mean an agent can replace a worker?
No. A benchmark measures performance on a defined task distribution. Jobs contain broader responsibilities, organizational knowledge, judgment, physical work, relationships and changing conditions that may not be represented in the dataset.
What is the Economic Replacement Ratio?
The Economic Replacement Ratio compares estimated competent-human cost with the total verified cost of achieving the same successful outcome using an agent workflow.
Should companies trust public agent benchmarks?
Public benchmarks are useful for comparison and shortlisting, but organizations should also evaluate candidate agents on private tasks drawn from their actual workflows before production deployment.
Will Decentralised News publish a leaderboard?
The framework is designed to support a public DN leaderboard. Scores should only be added after controlled, reproducible benchmark runs have been completed rather than estimated or inferred from vendor claims.
Disclosure
Decentralised News develops independent research frameworks, indices, benchmarks and decision tools covering AI, crypto and agentic finance. Some Decentralised News pages may contain affiliate or commercial relationships. These relationships do not determine benchmark inclusion, methodology, scores or editorial conclusions. The DN-RWAB framework is an original research methodology and does not constitute financial, employment or investment advice.
Related reading:
How Dangerous Is Your AI Agent? Measure Its Blast Radius
Coinbase’s Agentic Trading Stack: Equities, Crypto, x402 and Guardrails
Best AI Crypto Trading Tools for Beginners: 11 Platforms Compared
The Best AI Trading Demo Accounts for 2027
The AI Trading Authority Ladder: From Research Assistant to Autonomous Agent
Which Perp DEX Would You Trust an AI Agent to Trade On?
Best Crypto Platforms for AI Agents 2027 | Agentic Finance Rankings






