Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and…
Why am I seeing this Ranked on source trust — arXiv
It ranks mainly on source trust: arXiv is the most reliable outlet we track on this subject, and is the only one on the story so far.
Link-outLink-out, because it did not clear the bar for a write-up. Link-out means we point at the publisher and say nothing of our own.
| Factor | Weight | Score | Contribution | Where it came from |
|---|---|---|---|---|
| Corroboration | 0.35 | 0.39 | +0.135 26% | 1 independent org on the story. Tier-3 aggregators never corroborate — they can show something is circulating, never that it is true. |
| Source trustleads | 0.25 | 0.85 | +0.212 41% | arXiv is the highest-trust source on this story and is first-party — the organisation announcing its own news. Trust is taken from the best source, not averaged. |
| Pickup rate | 0.20 | 0.00 | +0.000 0% | One counted organisation, so there is no spread to measure — nothing has picked this up to set a rate. |
| Freshness | 0.20 | 0.85 | +0.171 33% | Halves every 10 hours from the newest item on the story. This is the only factor that rewards a story for nothing more than being recent. |
Corroboration counts distinct organisations, once each, and only from tiers 1 and 2. Freshness halves every 10 hours, so this ranking is a snapshot and will differ at the next build.

What happened
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories.
How this story arrived
Ordered by when each source was first observed, which is what the velocity figure is computed from. Publishers backdate; observed order does not.
- 01 Arxivfirst-party first seen Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory M
Overclock clusters coverage from independent sources and grades it automatically. The figures above are computed, not editorial. This page summarises and links to reporting by the outlets named — follow the links for the original work.