Skip to main content
Decentralised News Logo
AI Context Efficiency 2027: Useful Outcomes per Token
Agentic Finance

AI Context Efficiency 2027: Useful Outcomes per Token

By

Is your AI wasting context tokens? Use DN’s calculator to compare accepted outcomes, token consumption, costs and quality losses across matched agent workflows.

Decentralised News · Batch 2, Article 34

The Agent Context Efficiency Index 2027: Is Your AI Paying to Read the Same Information Again?

A bigger context window creates room for more information. It does not establish that every repeated token helps the agent finish the task.

By Heath Muchena · Published 2 October 2026 · DN-CEI v1.0 · 2027 planning edition

What Matters

Agent context efficiency measures accepted outcomes against all input tokens consumed across a workflow, including repeated history and retries. Smaller context is useful only when the task still succeeds. DN proposes reporting accepted tasks per 100,000 input tokens beside completion rate, cost and control failures. This edition provides a matched-workload method and calculator, without claiming live model rankings or measured savings.

DN Evidence Block

Last verified: 2 October 2026. Review period: primary research and engineering guidance reviewed on that date. DN benchmark sample: zero live model runs. Sources: the 2024 Lost in the Middle paper and Anthropic’s September 2025 context-engineering guidance. Author: Heath Muchena; no independent reviewer recorded.

Decisive distinction: published evidence supports testing how context is used; it does not establish a universal optimum length or current vendor ranking. DN-CEI is a proposed editorial metric. Calculator defaults are fictional, not research results.

Primary research · Primary engineering guidance

The DN Alpha Thesis: context has a carrying cost

An agent may request a small fact from a tool and then repeatedly carry that response through later model calls. Add the original instructions, previous outputs, tool descriptions and a growing conversation history, and the input consumed across the run can substantially exceed the unique information supplied by the user.

DN calls this the context carrying cost: the token exposure created by repeatedly making state available during a workflow. It is not automatically waste. Some repeated state preserves requirements and prevents mistakes. The research question is which information needs to remain active, which can be retrieved on demand and which can be replaced with a verified summary.

Measuring the largest individual prompt misses that question. A series of modest calls can consume more aggregate input than one large call. An apparent compression saving can disappear when the agent performs extra searches, retries or summary repairs.

This index therefore evaluates the entire declared workflow. It complements the Agent Model Efficiency Index, which focuses on successful work per dollar, and the memory benchmark, which tests retained state. Here the denominator is aggregate input-token consumption.

What the primary evidence tells us

Lost in the Middle evaluated multi-document question answering and key-value retrieval. In its studied settings, performance often depended on where relevant information appeared, with weaker use of information in the middle of the context. This historical result motivates position-sensitive tests; it does not prove identical behavior for every 2026 model or agent. [1]

Anthropic’s context-engineering guidance discusses curating instructions, tool information, retrieved data and history. It describes just-in-time retrieval, compaction and structured notes, while warning that aggressive compaction can discard important details. This is provider engineering guidance rather than an independent comparative benchmark. [2]

DN’s inference is practical: evaluate both removal and recovery. A compact representation is useful only if it preserves the information needed to meet the acceptance rule or provides a reliable way to retrieve it.

The DN Context Efficiency metric

MetricFormulaWhy it matters
Context efficiencyAccepted tasks ÷ aggregate input tokens × 100,000Connects token consumption to independently accepted work
Acceptance rateAccepted tasks ÷ attempted tasks × 100Exposes quality losses hidden by a smaller denominator
Input tokens per accepted taskAggregate input tokens ÷ accepted tasksIncludes failed attempts and repeated state in the workload
Cost per accepted taskTotal declared workflow cost ÷ accepted tasksTests whether token reduction translates into economic benefit
Matched-workload token reduction1 − candidate input tokens ÷ baseline input tokensDescribes consumption change, not proven avoidable waste

Sum model input usage across every call inside the system boundary. Include input occurrences served from cache, counting each occurrence once; do not add a cached subset twice if already included in the total. Include compression, routing and summary calls when they are part of the evaluated system. Do not count an output token again as output here; count its later occurrence if it becomes input to another call.

Internal processing that is not exposed in usage cannot be reconstructed from this metric. Record the provider’s usage semantics and the included components. Monetary cost should come from actual billing records or clearly labelled estimates, rather than assuming all input tokens share one price.

DN Context Waste Calculator

This diagnostic compares two versions of the same task set. It measures consumption change, not proven semantic waste. Defaults are fictional. Equal attempted counts are required, and the reader must verify that the underlying tasks match.

Baseline

Candidate

Baseline: 9.0 · Candidate: 17.6 accepted tasks / 100k input tokens

Use one currency for both cost inputs. Include the same cost categories, task identities and acceptance rules in both runs. The tool does not run models, inspect prompts or verify evidence. Inputs remain local to the page.

A fictional improvement that still needs investigation

Suppose the baseline attempts 100 tasks, accepts 90 and consumes one million input tokens at a total cost of 20 currency units. A candidate attempts the same tasks, accepts 88 and consumes 500,000 input tokens at a cost of 12.

Baseline context efficiency is 9 accepted tasks per 100,000 input tokens. Candidate efficiency is 17.6. Token consumption falls 50%, and cost per accepted task falls from about 0.2222 to 0.1364. But acceptance also falls from 90% to 88%.

The candidate is more efficient by the proposed token metric and less successful by the task acceptance measure. Investigate the lost outcomes before deployment. An omitted contractual requirement or permission boundary can be more consequential than the saving. None of these numbers is an observed model result.

Compare context strategies on the same contract

StrategyBest forAvoid ifCosts and accessPrimary failure risk
Full-history contextShort workflows where earlier details remain relevantHistory accumulates large irrelevant outputsRepeated input, latency and context-window constraintsRelevant evidence becomes hard to identify
Targeted retrievalLarge source collections with reliable identifiersRequired evidence cannot be located reliablySearch/index maintenance, tool calls and retrieval latencyMissing decisive information or retrieving stale material
Compacted historyLong runs with preservable decisions and task stateSummaries cannot retain required detailSummary calls, evaluation and recovery effortLoss of exceptions, constraints or provenance
Structured state and notesTasks with clear fields, decisions and unresolved itemsThe state schema leaves essential information outStorage, updates, versioning and validationStale or incorrectly written state

No provider is ranked, promoted or assigned an operational status in this edition. Costs depend on the actual implementation and service terms. For agentic finance, preserve authorization and transaction state outside an unverified narrative summary; the context strategy itself does not establish custody or execution authority.

A test pack that exposes information loss

Choose a fixed set of tasks and define success before running them. Use an independent evaluator and preserve the original evidence. Keep model, tokenizer, tools, decoding settings and task order controlled where possible. Change one context strategy at a time.

FixtureControlled changeAcceptance check
Position sensitivityMove a necessary fact among beginning, middle and end positionsThe correct answer remains grounded in the fact
Distractor loadAdd plausible but irrelevant recordsNo false selection or unsupported claim
Compaction boundarySummarize after introducing a rare exceptionThe exception survives and governs the final action
Stale evidenceProvide a dated value followed by an authoritative updateThe final result uses the right version
State recoveryClear history but preserve approved structured stateThe agent resumes without repeating a completed action
Authority preservationReduce context containing action restrictionsRestrictions remain enforced; no unauthorized action

Count all input consumed until final evaluation, not just the first request. Keep failed tasks in the denominator. Record median and tail latency separately; token reduction alone cannot establish faster delivery. Repeat enough runs to understand variability and publish sample counts rather than inventing confidence from a single clean demonstration.

What belongs in the evidence register

Each comparison needs task IDs, model and tokenizer versions, strategy configuration, source snapshots, per-call usage, cached-input treatment, total cost categories, acceptance outcomes and critical failures. Preserve the evaluator version and the exact system boundary.

Record how much context is submitted and which components are reintroduced: instructions, tool schemas, source evidence, tool results, history and structured state. That component ledger helps identify where to investigate. It does not prove that every repeated token is unnecessary.

Keep measurement and interpretation separate. “Input fell by 50%” can be a measured statement. “Half the context was waste” requires an additional causal argument showing that the removed information was unnecessary under the tested conditions.

Methodology, limitations and falsification

Method: review primary long-context research and provider engineering guidance; define accepted outcomes per aggregate input consumption; propose paired fixtures and quality gates; calculate diagnostics from reader-entered totals. No live model benchmark or vendor leaderboard has been performed.

Dataset status: the published fixture register is a proposed testing dataset, not observed performance data. The worked example is fictional. Future reports should publish sanitized task fixtures, per-call usage summaries, denominators and results on a permanent versioned URL.

Limits: tokenizers differ, task difficulty differs and stochastic outputs vary. Caching changes billing without necessarily changing logical input exposure. The metric does not measure semantic relevance, inaccessible internal compute or total infrastructure efficiency. Cross-model rankings need explicit caveats and additional outcome/cost measures.

Falsification test: reject a claim of useful context optimisation if matched evaluation finds lost necessary evidence, unacceptable quality decline, critical control failures or hidden auxiliary costs that reverse the claimed economic benefit.

Maintenance: rerun after model, tokenizer, tool, retrieval or compaction changes. Review sources quarterly. Change log: 2 October 2026, initial DN-CEI methodology and calculator. Corrections: send the claim, fixture and evidence through DN’s contact page.

Your next step: remove one source of repetition, then retest

Choose one repeated tool output or history segment. Test a targeted retrieval or verified summary alternative on a fixed task set. Compare token consumption, accepted outcomes, total cost and failure evidence together. Retain the original strategy until the comparison supports the change.

Connect workflow economics to the AI Agent ROI Index by Industry, and tool evidence to the MCP Server Reliability Index. Use DN Pathfinder when selecting a broader platform stack.

Sources

  1. Liu et al.: Lost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics, 2024.
  2. Anthropic: Effective context engineering for AI agents, 29 September 2025.

Frequently asked questions

What is agent context efficiency?

DN defines it as accepted task outcomes per 100,000 aggregate model-input tokens across an evaluated workload. Acceptance rate, cost, latency and critical failures must be reported separately.

Is this a measured model leaderboard?

No. This edition publishes a proposed test method and a calculator using reader inputs. DN has not run a comparative model or context-strategy benchmark.

Should cached input tokens be counted?

Yes for the input-token denominator, if reported in provider usage. Count each input occurrence once and include cached occurrences. Use actual invoices for monetary cost; cached tokens can have different billing treatment.

Is a shorter context always better?

No. Removing necessary evidence, constraints or unresolved state can reduce correctness. Compare outcomes on matched tasks before treating token reduction as improvement.

Does the calculator measure redundant tokens?

No. It measures aggregate token reduction, completion rates and cost per accepted task. It cannot identify which tokens were unnecessary or establish avoidable waste.

Do retries belong in the denominator?

Yes. Include all model calls, retries and auxiliary context-management calls within the declared system boundary, including work on failed tasks.

How do you compare different tokenizers?

Prefer comparisons within the same model and tokenizer. Across models, token counts are not an identical unit of text; report that limitation and compare accepted outcomes, cost and latency as well.

What blocks a claimed improvement?

A critical safety or authorization failure blocks DN’s proposed recommendation. A drop in acceptance rate requires investigation even if accepted outcomes per token increases.

Disclosure: no affiliate links, paid placements or vendor price claims appear. DN-CEI is proposed editorial methodology, not certification. Defaults are fictional and no live savings are claimed.

Get the most talked about stories directly in your inbox

Join the Decentralised News briefing for independent crypto, DeFi and AI analysis. No spam, unsubscribe anytime.