Policy Iteration with Human Feedback

Teaching agentic AI to learn expert reasoning

Rare-disease expertise is scarce. liteOdyssey turns expert review of model failures into an explicit, clinician-gated diagnostic policy that a reasoning model can execute, experts can inspect, and future models can reuse.

A written policy that survives contact with new diseases, new models, and clinical complexity.
59.3%Disease Recall@1 across 1,243 public benchmark cases
1,193Development-excluded public cases spanning 679 unseen diseases
515Undiagnosed Diseases Network patients in the clinical-cohort evaluation
Cross-model transfer

One policy. Four model backbones. Each improved.

The same PIHF-derived policy and biomedical tool workflow was executed unchanged across GPT-5.4, Qwen3.6-35B, and DeepSeek-V4 Flash and Pro. Each evaluated model improved over its matched parametric baseline.

59.3% liteOdyssey disease Recall@1 with the GPT-5.4 backbone
18.7–26.5% range across the four matched baseline LLMs
Open the full vector figure
Figure 2a: four parametric baseline language models achieve disease Recall at 1 from 18.7 to 26.5 percent, while liteOdyssey with the GPT-5.4 backbone achieves 59.3 percent.
Figure 2a · Pooled Disease Recall@1 across the same 1,243 public benchmark cases. Grey marks matched no-policy, no-tool baselines; blue marks liteOdyssey with the GPT-5.4 backbone.
Explore the full benchmark matrix Recall@1 and Recall@5 across LIRICAL and three PhenoPacket Store settings Panels a–e
Full Figure 2 benchmark matrix comparing parametric baselines in grey with liteOdyssey in blue across four execution models, at Recall at 1 and Recall at 5, for LIRICAL and mapped, unmapped, and extension PhenoPacket Store cases.

Reading the figureThe parametric baseline used the same language model and clinical features without the liteOdyssey policy or biomedical tools. The full system used the PIHF-derived policy-guided workflow with biomedical tool use. DeepSeek-family runs were scored with an independent model-family judge; exact-OMIM sensitivity analyses are reported in the manuscript supplement.

01 / THE METHOD

The policy improves. The model stays frozen.

PIHF adapts the evaluate-and-improve structure of policy iteration to a setting where the quality of diagnostic reasoning cannot be reduced to a single programmable reward.

Execute

Policy-guided rollout

A reasoning-capable model applies the current diagnostic workflow and records the evidence path it followed.

Evaluate

Stage-level critique

An expert-designed credit system identifies premature closure, weak evidence use, or missed corrective search.

Govern

Clinician gate

Experts judge whether a proposed revision is clinically sound and whether full-panel performance remains stable.

Consolidate

Versioned policy

Accepted feedback becomes a durable rule in a readable artifact, ready for the next model and the next case.

What liteOdyssey executes

An eight-phase clinical reasoning workflow

The workflow is structured but not linear. New evidence can send the agent back to reframe the pattern, widen the candidates, or challenge its leading diagnosis.

Pattern recognition
Candidate generation
Validity triage
Reasoning and ranking
Confidence assessment
Corrective search
Reflective adjudication
Ranked differential
02 / THE EVIDENCE

A reusable strategy, not a catalogue of remembered cases

The matched parametric baseline uses the same model and clinical features without the liteOdyssey policy or biomedical tools. Held-out analyses test whether the gain survives beyond policy-development cases.

Public benchmark union
59.3%

Disease Recall@1 across 1,243 cases and 722 rare diseases. Matched baseline: 26.5%.

LIRICAL
58.6%

Recall@1 versus 35.1% for the matched baseline.

PhenoPacket Store
59.6%

Recall@1 versus 22.9% for the matched baseline.

Development-excluded public evaluation
679 unseen diseases

The 1,193 held-out cases preserve nearly identical gains, arguing against case-specific policy adaptation.

Real-world clinical cohort
515 UDN patients

Recall@1 improved from 16.7% to 20.4%; blinded physician review independently corroborated the differential-diagnosis gain.

Transfer is the test.

The same natural-language policy improved closed- and open-weight reasoning models without changing their weights. The durable artifact is the procedure, not one model release.

03 / A PUBLIC TRACE

Watch the agent gather and weigh evidence

This recording shows the reviewer workspace running a full diagnostic walkthrough from a single prompt — phased reasoning, live tool use, and the final structured report. The live reviewer workspace below is interactive.

recorded reviewer workspace demonstration

Reviewer experience

Take the policy out for a real voyage.

Enter the protected workspace to question the system, inspect the reasoning process, try a phenotype profile, and stress-test where the policy holds or fails.

Request access