OVERCLOCK.news
ai

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

-cross Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream…

BlueskyXRedditMail
Revised 1 change recorded since we first saw this
Arxiv 2 versions
  1. Current Summary changed

    1 new sentence, beginning: “Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream…”

    When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

    Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test this proxy for pre-action policy classification in $\tau$-bench using Qwen2.5-3B/7B classifiers. Under gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by $0.20$ after tuning at 3B, with the same ordering at 7B under shared hyperparameters. We then replace the benchmark-designated governing rule with the top-ranked benchmark assertion retrieved from decision-time context. Although the exact governing rule is retrieved at rank 1 for only $7\%$ of airline states, the primary 3B classifier obtains macro-F1 $0.58$ with retrieved assertions versus $0.60$ with the gold rule ($\Delta=-0.02$, task-cluster 95\% CI $[-0.23,+0.21]$); random non-gold and no-assertion controls score $0.32$ and $0.21$. We do not detect a macro-F1 difference between retrieved assertions and the gold rule in this configuration, although the interval remains too wide to establish non-inferiority. The same qualitative pattern appears with a second retriever and at 7B, while varying across fine-tuning configurations. These results show that exact-match recovery of the benchmark-designated rule can underestimate the downstream utility of retrieved benchmark assertions in this setting. Retrieval should therefore be evaluated inside the classification loop rather than by exact-match recall alone.

  2. As first published

    When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

    -cross Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test this proxy for pre-action policy classification in $\tau$-bench using Qwen2.5-3B/7B classifiers. Under gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by $0.20$ after tuning at 3B, with the same ordering at 7B under shared hyperparameters. We then replace the benchmark-designated governing rule with the top-ranked benchmark assertion retrieved from decision-time context. Although the exact governing rule is retrieved at rank 1 for only $7\%$ of airline states, the primary 3B classifier obtains macro-F1 $0.58$ with retrieved assertions versus $0.60$ with the gold rule ($\Delta=-0.02$, task-cluster 95\% CI $[-0.23,+0.21]$); random non-gold and no-assertion controls score $0.32$ and $0.21$. We do not detect a macro-F1 difference between retrieved assertions and the gold rule in this configuration, although the interval remains too wide to establish non-inferiority. The same qualitative pattern appears with a second retriever and at 7B, while varying across fine-tuning configurations. These results show that exact-match recovery of the benchmark-designated rule can underestimate the downstream utility of retrieved benchmark assertions in this setting. Retrieval should therefore be evaluated inside the classification loop rather than by exact-match recall alone.

Read the current version at Arxiv →

Versions are compared on the headline and summary the publisher puts in their feed. An edit to the body of an article that leaves both untouched will not appear here.

Why am I seeing this Ranked on source trust — arXiv

It ranks mainly on source trust: arXiv is the most reliable outlet we track on this subject, and is the only one on the story so far.

The classifier could not identify the subject from the text, so the section was inherited from the source feed. We do not summarise what we cannot identify — this one links straight out.

Link-outLink-out, because the subject could not be identified from the text. Link-out means we point at the publisher and say nothing of our own.

Blended score 0.519 — every figure below is computed, none of it is editorial.
FactorWeightScore ContributionWhere it came from
Corroboration 0.35 0.39 +0.135 26% 1 independent org on the story. Tier-3 aggregators never corroborate — they can show something is circulating, never that it is true.
Source trustleads 0.25 0.85 +0.212 41% arXiv is the highest-trust source on this story and is first-party — the organisation announcing its own news. Trust is taken from the best source, not averaged.
Pickup rate 0.20 0.00 +0.000 0% One counted organisation, so there is no spread to measure — nothing has picked this up to set a rate.
Freshness 0.20 0.85 +0.171 33% Halves every 10 hours from the newest item on the story. This is the only factor that rewards a story for nothing more than being recent.

Corroboration counts distinct organisations, once each, and only from tiers 1 and 2. Freshness halves every 10 hours, so this ranking is a snapshot and will differ at the next build.

Read the full article at Arxiv →

What happened

-cross Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test this proxy for pre-action policy classification in $\tau$-bench using Qwen2.5-3B/7B classifiers. Under gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by $0.20$ after tuning at 3B, with the same ordering at 7B under shared hyperparameters. We then replace the benchmark-designated governing rule with the top-ranked benchmark assertion retrieved from decision-time context.

1independent orgs
52story score
0velocity
85source trust
1passes seen

How this story arrived

Ordered by when each source was first observed, which is what the velocity figure is computed from. Publishers backdate; observed order does not.

  1. 01 Arxivfirst-party first seen When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

Overclock clusters coverage from independent sources and grades it automatically. The figures above are computed, not editorial. This page summarises and links to reporting by the outlets named — follow the links for the original work.