arXiv 2607.13172v1Jul 14, 2026

Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

Ilias Kazantzidis et al.

Brief context

Publication timing, weekly edition context, and source links for this brief.

Published

Jul 14, 2026, 6:22 PM

Current score

72

Original paper

The executive brief below is grounded in the source paper and linked back to the arXiv abstract.

We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown and no suitable reward function is available. In the context of safety-critical environments, we consider traditional reinforcement learning impractical and resort to the resource of human input. We introduce DROPJ, a human-centred method for both safe training and deployment. We first learn a world model (a learned simulator) from a dataset of prior real-world trajectories. A human then plays the game in this learned simulator to extract several informative simulated trajectories. From these, we sample pairs of simulated trajectory segments and elicit from a human their preference over these segments, as well as a reason (justification) for their choice. We then train a reward model from these justified preferences and use it, together with the world model, to directly deploy the agent using model predictive control. Running real-user experiments, we find that generating informative simulated trajectories from a user significantly reduces the computational cost during training compared to other strategies, and can also improve the performance during deployment. In the context of training within a learned simulator, we show that the use of preferences rather than other types of feedback substantially improves the performance during deployment. We further demonstrate that safety justifications accompanying preferences can significantly enhance safety or prioritise user-prescribed aspects of safety associated with them during deployment.

Score 72Full-paper briefagentstrainingmodelsdata

Executive brief

A short business-reader brief that explains why the paper matters now and what to watch or do next.

Why this is worth your attention

DROPJ is a credible attempt to move safe-agent training away from dangerous trial-and-error and brittle hand-written reward rules. The paper’s practical claim is that a learned simulator plus short bursts of expert preference feedback can make safety-critical agent training cheaper in human time and safer during training, while justifications let teams specify which risks matter most. The catch is material: the results are still in simulated driving-style tasks, and the whole approach depends on a good world model built from sufficiently broad real trajectories.

  • The business-relevant move is not a better racing-game agent; it is a workflow for training agents where no one can write a complete reward function and unsafe trial-and-error is unacceptable. If it generalizes, operations teams could convert logged experience plus expert judgment into deployable control policies without hand-coding every objective.
  • The paper’s strongest operational claim is that expert input can be concentrated into short simulator sessions and preference judgments: one setup used 10 simulated episodes at roughly 30 seconds each, while preference methods beat sparse labels after about 20 minutes of feedback. Watch for this pattern in vendor demos: less time waiting on generated queries, more time giving high-value judgments.
  • A learned simulator is only useful if it has seen enough of the real operating envelope, including failures or controlled unsafe cases. For buyers, the hard question is not “do you use a world model?” but “what evidence shows your model covers rare, risky, and edge-case states well enough to plan against them?”
  • The paper treats human explanations as tunable safety priorities, not just training labels. That is strategically interesting because it points toward systems where risk appetite can be encoded explicitly, but the authors also show interacting risks can create carry-over effects, so this is not a simple compliance dial.
  • The evidence is real-user empirical work, but still in simulated car-racing environments, and the paper itself reports hallucinated simulator frames, skipped queries, and a safety-performance trade-off. The right takeaway is a promising training pattern, not proof that world-model agents are ready for unbounded physical deployment.

Evidence ledger

The strongest claims in the brief, along with the confidence and citation depth behind them.

capabilityhighp.3p.4

DROPJ proposes a practical pipeline for training safe agent behavior from offline trajectories, simulated expert interaction, preferences, justifications, and model-predictive control.

trainingmediump.20p.21

Preference feedback, especially with justifications, appears more human-time-efficient than sparse labels in the reported experiments.

traininghighp.18p.18

Training inside a learned simulator avoids real-environment unsafe training episodes in the reported setup, unlike model-free RL trained directly in the environment.

caveathighp.16p.34

World-model quality and cost remain material bottlenecks, both computationally and operationally.

Related briefs

More plain-English summaries from the archive with nearby topics or operator relevance.

cs.SE

TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution

Jiale Amber Wang, Kaiyuan Wang, Pengyu Nie

cs.CL

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

A. Sayyad et al.

cs.LG

Adaptive Inference Batching using Policy Gradients

Ruslan Sharifullin

cs.SE

Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs

Sheng-Wei Peng, Yi-Hsun Lin, Yi-Pei Lee

Thank you to arXiv for use of its open access interoperability. This product was not reviewed or approved by, nor does it necessarily express or reflect the policies or opinions of, arXiv.
LightDark