Skip to main content
Decentralised News Logo
The 100-Task AI Agent Test: Which Agents Can Actually Do Useful Work?
Agentic Finance

The 100-Task AI Agent Test: Which Agents Can Actually Do Useful Work?

By

The DN Real-World Agent Benchmark tests AI agents on 100 tasks humans actually pay for, measuring completion, correctness, cost, reliability and human intervention.

Decentralised News Research · Agentic Finance

The Real-World Agent Benchmark: 100 Tasks Humans Actually Pay For

AI agents are becoming impressive at benchmarks. But businesses do not pay agents for benchmark points. They pay for completed research, reconciled accounts, qualified leads, fixed software, resolved support tickets and other useful outcomes. The DN Real-World Agent Benchmark proposes a harder test: can an agent reliably finish work someone would actually pay a human to do?

Framework: DN-RWAB 1.0 · Last verified: 29 September 2026 · Dataset design: 100 economically useful tasks across 10 work categories

What Matters

The next useful AI-agent leaderboard should not ask only whether an agent can solve a coding issue, browse a sandbox or call the correct API. It should ask whether the agent can deliver paid-quality work. DN-RWAB scores 100 practical tasks on completion, correctness, human intervention, time, cost, reliability and economic value.

DN Evidence Block

100 paid-work tasks in the proposed dataset
10 economic work categories
7 scoring dimensions per task
3× minimum repeat runs proposed per task

Why this exists: Current agent benchmarks provide valuable evidence about coding, browser use, computer control, tool use and task horizon. But these tests do not by themselves establish whether an agent can repeatedly perform the heterogeneous work businesses and individuals buy in the open economy.

The Signal: Benchmark Scores Are Not the Same as Economic Utility

AI benchmarking has progressed rapidly.

SWE-bench evaluates whether systems can solve genuine software-engineering issues. Its Verified dataset contains 500 tasks reviewed by humans for clarity and solvability.

OSWorld moves evaluation into real computer environments rather than static prompts. Its newer generations increasingly emphasize longer, more realistic application workflows.

Princeton's τ-bench family evaluates agents that must converse with users, invoke tools and follow policies inside realistic domains such as airline support.

METR measures task-completion time horizons by comparing agent success against the amount of time expert humans require to complete tasks.

These are meaningful advances.

But each benchmark measures a slice of useful work.

DN thesis: The economically decisive benchmark is not “Can an agent perform this benchmark task?” It is “Can the agent deliver an outcome that a real customer, employer or business would otherwise pay someone to produce?”

The Missing Unit: Paid-Quality Completion

Work has several characteristics that ordinary benchmark success rates can hide.

A research report can contain the requested facts and still be unusable because its sources are wrong.

A sales prospect list can contain 100 rows and still be worthless because half the contacts are irrelevant.

A spreadsheet can calculate correctly while silently corrupting the formatting or formulas elsewhere.

A customer-service agent can eventually reach the correct outcome but consume more time and money than a trained employee.

An agent can successfully finish a task once and fail the next four times.

That means binary completion is insufficient.

DN proposes a higher standard:

Paid-Quality Completion: a task is successful only when the output is complete, materially correct, usable without hidden repairs and delivered at a cost and reliability level that could plausibly substitute for paid human work.

The DN 100-Task Real-World Agent Benchmark

DN-RWAB 1.0 divides 100 tasks across ten categories representing commercially meaningful knowledge work.

Category Tasks Example paid work Primary failure mode
Research & Analysis 10 Competitive research, source verification, industry briefings Unsupported conclusions or missed evidence
Software & Technical Work 10 Bug fixes, scripts, debugging, data transformations Correct-looking output that breaks elsewhere
Finance & Accounting 10 Reconciliation, categorization, variance analysis, invoice checks Silent numerical or classification errors
Sales & Lead Generation 10 Prospect research, qualification, outreach preparation Volume without relevance or accuracy
Marketing & Content 10 Campaign research, briefs, SEO analysis, publishing workflows Plausible but generic output
Customer Operations 10 Support resolution, refunds, account changes, escalation Policy violations or incorrect actions
Commerce & Procurement 10 Product sourcing, quote comparison, purchasing decisions Wrong specifications or hidden total cost
Administrative Work 10 Scheduling, document processing, inbox triage, travel planning Missed constraints or incomplete execution
Data & Spreadsheet Work 10 Cleaning, analysis, formulas, reports, dashboard preparation Subtle data corruption
Agentic Finance 10 Route comparison, treasury checks, transaction preparation, risk monitoring Financial loss or policy breach

What the 100 Tasks Should Actually Look Like

A useful benchmark should avoid toy requests that can be solved by pattern matching.

Each DN task should contain realistic constraints, imperfect information and a verifiable outcome.

1

Research

Identify five operational competitors to a specified company, verify their current offerings from primary sources and deliver a comparison with dated evidence.

2

Spreadsheet

Clean a messy transaction export, preserve source data, categorize expenses and produce a reconciled summary without breaking formulas.

3

Customer Support

Resolve a customer's account problem while following refund, privacy and escalation policies.

4

Procurement

Compare suppliers for an exact specification while incorporating delivery, minimum order quantities and total landed cost.

5

Software

Diagnose a real failure, implement a fix, verify regressions and explain the change in language another developer could review.

6

Finance

Reconcile payments against invoices and flag exceptions that require human investigation rather than inventing explanations.

DN's Seven-Dimensional Scoring Model

Every benchmark run should generate more than a pass or fail.

Metric Weight What it measures
Task completion 25% Whether the requested outcome was actually achieved
Correctness 20% Whether facts, calculations and actions were materially accurate
Output usability 15% Whether the deliverable could be used without substantial repair
Reliability 15% Whether repeated runs produce consistently successful outcomes
Human intervention 10% How much rescue, clarification or correction was required
Cost efficiency 10% Agent cost compared with the estimated human cost of completing the same task
Time efficiency 5% Elapsed completion time relative to a competent human baseline

Try the DN Economic Task Score

DN Economic Task Calculator

Estimate whether an agent output qualifies as economically useful work rather than merely technical completion.

Was the actual requested outcome achieved?
How accurate was the finished work?
Could a customer or employer actually use it?
How consistently did repeated runs succeed?
How much human rescue was necessary?
How does agent cost compare with human completion cost?
How quickly did the agent finish?
DN Economic Task Score
0/100

Not Yet Economically Useful

Select the observed result for each dimension.

Interpretation
  • Measure the finished outcome, not the agent's reasoning style.
  • Repeat the task before drawing production conclusions.

DN Economic Task Score Bands

Score Classification Meaning
0–39 Not Economically Useful The agent may demonstrate capability but cannot yet substitute for paid work.
40–59 Human-Dependent Useful as assistance, but human correction materially contributes to the final outcome.
60–79 Commercially Useful The agent can complete meaningful work but still needs defined review and failure handling.
80–100 Paid-Quality Autonomous The agent repeatedly produces usable results at commercially compelling cost and reliability.

The Benchmark Must Penalize Lucky Runs

One successful demonstration is not reliability.

Agent systems are probabilistic, and multi-step workflows create more opportunities for small differences to compound.

DN-RWAB therefore proposes at least three independent runs per task for baseline testing.

High-value or safety-sensitive tasks should eventually use more.

DN Repeatability Rule: An agent does not “pass” a paid-work task merely because it succeeded once. Commercial usefulness requires repeatable outcomes.

Why Repeatability Changes the Leaderboard

Imagine two agents.

Agent A successfully completes 85 of 100 tasks on its best run.

Agent B completes 78.

That headline makes Agent A appear better.

Now run each task five times.

If Agent A succeeds unpredictably while Agent B produces the same correct result almost every time, the commercial conclusion changes.

Businesses often care more about predictable 78% performance than an unstable system capable of occasionally reaching 85%.

Public Benchmarks Already Show Why Harnesses Matter

Agent evaluation measures more than the underlying model.

The same model can perform differently depending on its tools, prompts, memory architecture, environment, retry strategy and agent scaffolding.

SWE-bench explicitly distinguishes model comparisons from broader agent-system submissions. Princeton has also removed compromised benchmark results when test-set leakage was identified in a τ-bench scaffold.

That is important.

DN rule: The benchmark entry should identify the entire tested system, not imply that the result measures the underlying model alone.

DN-RWAB Should Score the System, Not the Brand

Every leaderboard entry should record:

  • model and exact version;
  • agent framework or harness;
  • tools available;
  • maximum steps;
  • reasoning or compute configuration where disclosed;
  • external search access;
  • memory configuration;
  • human interventions;
  • retries;
  • total tokens or compute usage where measurable;
  • API and tool cost;
  • wall-clock completion time;
  • test date;
  • task version.

Without these details, a leaderboard risks comparing fundamentally different systems under one model name.

The Cost per Successful Task Is More Useful Than Token Price

Cheap tokens do not necessarily create cheap work.

Suppose Agent A costs $0.20 per attempt but succeeds only 30% of the time. Its effective model cost per successful task is approximately $0.67 before human rescue costs.

Agent B might cost $0.50 per attempt but succeed 90% of the time, producing an effective cost near $0.56 per successful completion.

The supposedly more expensive agent is economically better.

This is why DN proposes:

Cost per Successful Task = Total Agent + Tool + Human-Intervention Cost ÷ Verified Successful Tasks

Human Intervention Must Be Priced

One of the largest hidden costs in agent deployments is human rescue.

An agent may cost pennies to run but require ten minutes of professional review after every task.

At scale, the review cost can dominate inference.

DN-RWAB therefore treats human involvement as part of the economic cost rather than pretending it is free.

Intervention level Example DN treatment
None Agent completes and verifies task independently 0 rescue minutes
Light review Human checks final output without materially changing it Record review time
Correction Human fixes errors before output becomes usable Deduct usability and intervention points
Rescue Human must complete a failed part of the workflow Material score penalty
Takeover Human finishes the task Agent does not receive full completion credit

The Economic Replacement Ratio

DN proposes another metric for the benchmark:

Economic Replacement Ratio = Estimated Human Cost ÷ Total Verified Agent Completion Cost

An Economic Replacement Ratio of 1 means agent and human cost are approximately equivalent.

A ratio of 5 means the verified agent workflow costs roughly one-fifth of the competent human baseline.

But cost advantage alone is insufficient.

A cheap workflow producing unusable work has no meaningful replacement value.

Example: A $100 Human Task vs a $5 Agent Run

Consider a market-research task that a competent freelancer would charge approximately $100 to complete.

The agent workflow costs $5.

At first glance, the economic replacement ratio appears to be 20×.

But suppose the agent produces a usable answer only half the time and requires 20 minutes of human checking and correction.

The genuine cost advantage could be dramatically lower.

Benchmarking therefore has to measure the entire production loop.

Ten Research Tasks Humans Actually Pay For

# Task Verified outcome
1Competitor landscapeCurrent competitors, evidence and differentiated comparison
2Supplier researchQualified suppliers meeting exact constraints
3Market sizingTransparent assumptions and reproducible calculation
4Regulatory researchCurrent primary-source requirements and jurisdiction boundaries
5Company diligenceVerified facts separated from inference
6Literature reviewRelevant papers, findings and limitations
7Product comparisonCurrent features, pricing and decision-relevant differences
8News intelligenceMaterial developments with publication and event dates distinguished
9Executive briefingConcise decision-ready synthesis
10Evidence auditUnsupported claims identified and corrected

Ten Finance and Accounting Tasks

# Task Verified outcome
11Bank reconciliationTransactions reconciled with exceptions isolated
12Invoice matchingCorrect invoice-payment mapping
13Expense categorizationConsistent categories with ambiguous items flagged
14Budget variance analysisCorrect variances and material drivers identified
15Cash-flow forecastTransparent assumptions and internally consistent projection
16Duplicate-payment checkTrue duplicates isolated without excessive false positives
17Management reportingCorrect financial summary built from supplied data
18Collections prioritizationReceivables correctly segmented by urgency
19Purchase-order checkPO, invoice and receipt discrepancies identified
20Financial anomaly reviewMaterial anomalies flagged without invented explanations

Ten Sales Tasks

# Task Verified outcome
21ICP company sourcingCompanies genuinely matching defined criteria
22Lead qualificationRelevant prospects prioritized accurately
23Decision-maker researchCorrect role and current employment verified
24Account researchUseful trigger events and pain points identified
25Meeting preparationAccurate briefing with no fabricated personal details
26CRM cleanupDuplicates and malformed records corrected safely
27Opportunity prioritizationPipeline ordered using stated evidence
28Proposal researchProposal tailored to verified buyer requirements
29Follow-up draftingRelevant follow-up grounded in actual interaction
30Lost-deal analysisPatterns separated from speculation

Ten Marketing and Content Tasks

# Task Verified outcome
31SEO competitor analysisReal content gaps supported by search evidence
32Content brief creationPublishable brief aligned with user intent
33Source-backed articleAccurate claims with current sources
34Campaign analysisCorrect performance interpretation from supplied metrics
35Social repurposingChannel-appropriate derivatives preserving meaning
36Newsletter productionAccurate concise edition built from source material
37Landing-page auditSpecific conversion issues identified rather than generic advice
38Editorial fact checkUnsupported or stale claims surfaced
39Keyword clusteringSearch terms grouped by meaningful intent
40Content refreshStale claims updated without degrading existing useful material

Ten Software and Technical Tasks

#TaskVerified outcome
41Bug diagnosisRoot cause correctly identified
42Bug repairIssue fixed without regression
43Data-conversion scriptCorrect output across test cases
44API integrationWorking implementation with failures handled
45Test creationTests catch intended failure modes
46Documentation updateDocumentation accurately matches implementation
47Dependency migrationMigration completed without hidden breakage
48Performance diagnosisBottleneck supported by evidence
49Security configuration reviewMaterial misconfigurations correctly identified
50Deployment troubleshootingProduction issue resolved and verified

Ten Customer Operations Tasks

#TaskVerified outcome
51Order-status resolutionCorrect current status communicated
52Refund processingPolicy-compliant refund outcome
53Account-access issueProblem solved without bypassing security controls
54Subscription changeCorrect plan change executed
55Complaint handlingIssue resolved or appropriately escalated
56Billing disputeEvidence checked before action
57Technical support triageCorrect category and next action
58Policy interpretationPolicy applied consistently
59Case summarizationComplete accurate handoff note
60Escalation decisionHigh-risk case escalated at correct threshold

Ten Commerce and Procurement Tasks

#TaskVerified outcome
61Product sourcingProducts satisfy all mandatory requirements
62Quote normalizationSupplier quotes converted to comparable terms
63Total-cost comparisonFreight, taxes, minimums and extras included
64Vendor verificationSupplier existence and claims checked
65Purchase recommendationRecommendation follows stated constraints
66Inventory exceptionShortage correctly identified and routed
67Contract comparisonMaterial commercial differences extracted accurately
68Renewal reviewPricing and contractual changes surfaced
69Alternative supplier searchViable substitutes meeting specification
70Purchase-order preparationCorrect quantities, pricing and supplier details

Ten Administrative Tasks

#TaskVerified outcome
71Calendar coordinationMeeting scheduled without conflict
72Travel planningItinerary respects time, price and travel constraints
73Inbox triageImportant messages prioritized correctly
74Document organizationFiles classified without loss or duplication
75Meeting-note extractionDecisions and actions captured correctly
76Form completionFields correctly populated from evidence
77Deadline trackingDates and dependencies correctly identified
78Expense submissionReceipts matched and policy rules respected
79Contact-data cleanupRecords normalized without merging distinct people
80Action follow-upOutstanding commitments accurately surfaced

Ten Data and Spreadsheet Tasks

#TaskVerified outcome
81Dataset cleaningErrors corrected without deleting valid variation
82Formula repairBroken formulas fixed consistently
83Pivot analysisCorrect aggregation and grouping
84Dashboard preparationMetrics correspond exactly to source data
85Duplicate detectionTrue duplicates isolated with low false positives
86Outlier reviewPotential anomalies identified without automatic deletion
87Data mergeDatasets joined without record corruption
88CSV transformationTarget schema produced accurately
89Metric calculationFormula and final value both verifiable
90Trend analysisConclusions supported by underlying data

Ten Agentic Finance Tasks

#TaskVerified outcome
91Execution-route comparisonBest route identified using real total execution cost
92Stablecoin transfer planningNetwork, fees and destination compatibility verified
93Treasury-balance reviewBalances and exposures correctly summarized
94Wallet-permission auditMaterial permissions and risks identified
95Transaction simulationExpected state change correctly represented
96Yield comparisonNet yield separated from headline APY and risk
97Protocol-risk reviewCurrent material risks supported by evidence
98Position monitoringThreshold breach correctly detected
99ReconciliationOn-chain activity reconciled to internal records
100Controlled transaction preparationCorrect transaction prepared within explicit policy limits

The Hardest Part Is Ground Truth

Real-world tasks are more difficult to benchmark than multiple-choice questions because the correct result may contain judgment.

DN-RWAB should therefore use three types of evaluation.

Deterministic

Use exact checks wherever possible: balances, formulas, successful transactions, code tests, record states and database outputs.

Rubric Based

Use explicit criteria for outputs such as research briefs where several valid answers may exist.

Expert Reviewed

Use blinded human reviewers for tasks requiring professional judgment, with disagreements documented.

The Human Baseline Matters

An agent cannot be described as economically competitive without a comparison point.

Every task should therefore estimate:

  • competent human completion time;
  • competent human cost;
  • expected human error range where measurable;
  • minimum acceptable professional quality.

This does not mean humans must achieve 100%.

In fact, comparing agents against realistic human performance rather than an imaginary perfect worker makes the benchmark more useful.

Task Horizon Is Becoming a Critical Metric

METR's task-horizon work provides an important insight: agent capability is not only about whether a system can do something, but how long and complex a human-equivalent task it can complete reliably.

Real employment contains tasks ranging from two-minute lookups to projects that consume days.

DN-RWAB should therefore record estimated competent-human completion time for every task.

Human task duration DN class Examples
Under 15 minutesMicroLookup, categorization, simple update
15–60 minutesShortResearch note, reconciliation, lead qualification
1–4 hoursProfessionalDetailed analysis, software repair, data workflow
4–8 hoursWorkdayMajor report, system migration, complex diligence
8+ hoursLong HorizonMulti-stage professional project

Why OSWorld 2.0 Matters to This Thesis

The direction of recent benchmark design already points toward longer real-world workflows.

OSWorld 2.0 introduced 108 long-horizon computer-use workflows whose human completion time had a median of roughly 1.6 hours.

That is materially closer to professional work than a short GUI task.

Its published research still found frontier systems struggling with the full workflows, especially when success required tracking hidden state, changing information and multiple constraints.

The conclusion is not that agents are incapable.

It is that realistic work exposes failure modes that simpler tasks can miss.

The Benchmark Should Include Failure Recovery

Humans encounter broken links, missing data, failed APIs, ambiguous requests and changed circumstances.

Real agents will too.

A practical benchmark should therefore deliberately include recoverable failures.

Examples:

  • a supplier page becomes unavailable;
  • a spreadsheet contains an unexpected column;
  • a user provides conflicting instructions;
  • a tool call fails temporarily;
  • a booking option disappears;
  • a transaction simulation fails;
  • a source contradicts another source;
  • required information is missing.

A robust agent should know whether to retry, use another route, ask for clarification, escalate or stop.

Real-world intelligence includes knowing when not to proceed.

The Benchmark Should Penalize Confident Guessing

One particularly expensive agent failure is plausible fabrication.

An agent that asks for missing information may seem less autonomous than one that invents an answer and proceeds.

Economically, the cautious agent may be far more useful.

DN-RWAB should therefore reward appropriate abstention.

If success requires information the system does not possess, the correct behavior may be to stop and request it.

Agents Should Be Measured on Net Labor Saved

Gross automation can be misleading.

A workflow may eliminate 60 minutes of manual work while creating 25 minutes of checking, correction and system maintenance.

The real productivity gain is smaller.

DN Net Labor Saved = Human Baseline Time − Agent Run Oversight − Correction Time − Required Human Completion Time

This makes it possible to compare systems that automate different portions of the same workflow.

The DN Agent Economic Frontier

Over time, the most revealing visualization may not be a single ranking.

It may be a frontier.

Plot each tested system by:

  • paid-quality completion rate;
  • cost per successful task;
  • human intervention minutes;
  • median task horizon;
  • repeatability;
  • failure severity.

Different systems may dominate different parts of this frontier.

One agent might excel at short low-cost administrative work.

Another may be expensive but capable of completing long technical projects.

A third may be exceptionally reliable in financial workflows where correctness matters more than raw speed.

Why a Single “Best AI Agent” Is Usually the Wrong Question

Agent performance is task dependent.

A coding agent should not automatically be considered superior to a customer-service agent because it performs better on a software benchmark.

DN-RWAB should therefore publish:

  • overall score;
  • category scores;
  • task-horizon distribution;
  • cost per successful task;
  • reliability;
  • human-intervention requirement.

Readers can then select systems according to the work they actually want performed.

Proposed DN Leaderboard Structure

Agent Paid-Quality Completion Reliability Median Cost / Success Human Minutes Longest Reliable Task
Agent A Awaiting benchmark run — — — —
Agent B Awaiting benchmark run — — — —
Agent C Awaiting benchmark run — — — —

DN will not fabricate benchmark scores before controlled runs are completed. Initial leaderboard cells should remain explicitly unscored until reproducible test data exists.

How DN-RWAB Differs From Existing Agent Benchmarks

Benchmark Primary question DN-RWAB adds
SWE-bench Can the system fix real software issues? Cross-industry paid work, cost and human-intervention accounting
OSWorld Can the agent operate real computer interfaces? Economic usefulness across professions and channels
τ-bench Can an agent interact with users and tools while following domain rules? Broader work categories and explicit labor economics
METR Time Horizon How long a human-equivalent task can an agent complete at a given reliability? Task value, output usability and cost per successful outcome
DN-RWAB Can this system reliably produce work someone would pay for? Economic completion as the central unit

For Individuals: Measure What You Can Stop Paying Someone Else to Do

For freelancers, creators and professionals, the useful question is not how many benchmarks a model tops.

It is which tasks can now be delegated without increasing mistakes.

Start by documenting recurring paid or time-consuming work:

  • research;
  • data cleaning;
  • report preparation;
  • prospecting;
  • bookkeeping preparation;
  • content repurposing;
  • administration.

Test the agent repeatedly and compare the net labor saved.

For Businesses: Build a Private Version of the Benchmark

Public benchmarks should be treated as evidence, not production guarantees.

A company deploying agents should construct a private evaluation set from its own recurring work.

Twenty authentic internal tasks may reveal more about deployment readiness than hundreds of unrelated public benchmark tasks.

The strongest approach is therefore:

  1. use public benchmarks to shortlist systems;
  2. use DN-RWAB-style economic metrics to compare commercial usefulness;
  3. use private internal tasks before granting production authority;
  4. continue measuring live outcomes after deployment.

For Agent Developers: Optimize for Completed Economic Work

DN-RWAB creates a different engineering target.

Instead of optimizing only for tokens, benchmark percentage or tool-call accuracy, developers would optimize for:

  • correct outcome per dollar;
  • completed work per hour;
  • low human rescue;
  • high repeatability;
  • safe abstention;
  • reliable recovery;
  • professional output quality.

Those metrics map more directly to adoption.

The Economic Threshold for Agent Adoption

An agent becomes commercially interesting when four conditions intersect:

1. Capability

The agent can genuinely perform the target workflow.

2. Reliability

It performs well repeatedly rather than occasionally.

3. Economics

Total cost is competitive after oversight and corrections.

4. Risk

The consequences of failure are acceptable and controlled.

Improving model intelligence addresses only the first condition.

The commercial agent economy depends on all four.

DN Methodology

Framework: DN Real-World Agent Benchmark 1.0

Objective: Measure whether autonomous or semi-autonomous AI systems can repeatedly produce economically useful outcomes comparable with work humans are paid to perform.

Proposed dataset: 100 tasks across ten work categories.

Primary metrics:

  • 25% task completion;
  • 20% correctness;
  • 15% output usability;
  • 15% reliability;
  • 10% human intervention;
  • 10% cost efficiency;
  • 5% time efficiency.

Repeat testing: Minimum three independent runs per system-task combination in the first dataset release, with additional repetitions preferred for high-variance tasks.

Human baselines: Each task should include estimated competent-human completion time and cost, based on documented professional rates or commissioned human runs where practical.

System identification: Scores apply to the tested model, scaffold, tools, configuration and environment together.

Evidence boundary: The 100-task framework in this article is a benchmark specification. DN is not presenting invented performance results. Leaderboard rankings should be published only after controlled runs are completed.

Update cadence: Dataset versions should be permanent and reproducible. New task releases should use version numbers rather than silently changing historical tests.

Commercial independence: Sponsorship or affiliate availability must not change task selection, benchmark scoring or published results.

Falsification test: DN-RWAB should be revised if its score fails to correlate with independently measured real-world usefulness, economic savings or successful production deployment.

Limitations

No 100-task benchmark can represent the full labor market.

Work quality can be subjective, human cost varies by geography, and some jobs depend on tacit knowledge, relationships or physical actions that digital agents cannot reproduce.

Benchmark systems can also improve specifically against known datasets.

DN should therefore rotate private holdout tasks, publish methodology transparently and avoid presenting one aggregate score as universal intelligence.

Primary Sources and Benchmark References

SWE-bench Verified
Human-validated real-world software-engineering tasks used to evaluate coding agents.
SWE-bench Verified
OSWorld
Benchmarking multimodal agents inside real computer environments, including the newer long-horizon OSWorld 2.0 direction.
OSWorld
Princeton Holistic Agent Leaderboard / τ-bench
Evaluation of realistic tool-agent-user interactions and agent reliability.
τ-bench Airline
METR Task-Completion Time Horizons
Measurement of the human-equivalent task duration frontier AI agents can complete at specified reliability levels.
METR Time Horizons
OpenAI Economic Research
Evidence on the growing use of agents for longer-horizon workplace tasks and delegated work.
How Agents Are Transforming Work

Find the Right Agentic Stack

Use DN Pathfinder to compare the platforms and infrastructure relevant to what you actually want an agent to accomplish.

Open DN Pathfinder

Frequently Asked Questions

What is the DN Real-World Agent Benchmark?

The DN Real-World Agent Benchmark is a proposed 100-task evaluation framework designed to measure whether AI agents can repeatedly complete economically useful work that humans are currently paid to perform.

How is DN-RWAB different from other AI agent benchmarks?

Existing benchmarks often focus on a particular capability such as coding, browser use, tool calling or computer control. DN-RWAB adds cross-industry tasks and explicitly measures cost, human intervention, repeatability and paid-quality output.

What does paid-quality completion mean?

Paid-quality completion means that an agent has not merely technically finished a task. Its output must be materially correct, usable without substantial repair and sufficiently reliable that someone could plausibly pay for the result.

Why does the benchmark use repeated runs?

AI-agent behavior can vary between runs. Repetition helps distinguish a reliable system from one that occasionally succeeds by chance.

What is cost per successful task?

Cost per successful task divides total agent, tool and human-intervention costs by the number of verified successful outcomes. It is often more commercially meaningful than token price alone.

Why include human intervention in agent costs?

Human review, corrections and rescue work consume labor. Ignoring this can make an apparently cheap autonomous workflow look substantially more economical than it really is.

Does a high benchmark score mean an agent can replace a worker?

No. A benchmark measures performance on a defined task distribution. Jobs contain broader responsibilities, organizational knowledge, judgment, physical work, relationships and changing conditions that may not be represented in the dataset.

What is the Economic Replacement Ratio?

The Economic Replacement Ratio compares estimated competent-human cost with the total verified cost of achieving the same successful outcome using an agent workflow.

Should companies trust public agent benchmarks?

Public benchmarks are useful for comparison and shortlisting, but organizations should also evaluate candidate agents on private tasks drawn from their actual workflows before production deployment.

Will Decentralised News publish a leaderboard?

The framework is designed to support a public DN leaderboard. Scores should only be added after controlled, reproducible benchmark runs have been completed rather than estimated or inferred from vendor claims.

Disclosure

Decentralised News develops independent research frameworks, indices, benchmarks and decision tools covering AI, crypto and agentic finance. Some Decentralised News pages may contain affiliate or commercial relationships. These relationships do not determine benchmark inclusion, methodology, scores or editorial conclusions. The DN-RWAB framework is an original research methodology and does not constitute financial, employment or investment advice.

Get the most talked about stories directly in your inbox

Join the Decentralised News briefing for independent crypto, DeFi and AI analysis. No spam, unsubscribe anytime.