Policy-guided rollout
A reasoning-capable model applies the current diagnostic workflow and records the evidence path it followed.
Policy Iteration with Human Feedback
Rare-disease expertise is scarce. liteOdyssey turns expert review of model failures into an explicit, clinician-gated diagnostic policy that a reasoning model can execute, experts can inspect, and future models can reuse.
The same PIHF-derived policy and biomedical tool workflow was executed unchanged across GPT-5.4, Qwen3.6-35B, and DeepSeek-V4 Flash and Pro. Each evaluated model improved over its matched parametric baseline.
Reading the figureThe parametric baseline used the same language model and clinical features without the liteOdyssey policy or biomedical tools. The full system used the PIHF-derived policy-guided workflow with biomedical tool use. DeepSeek-family runs were scored with an independent model-family judge; exact-OMIM sensitivity analyses are reported in the manuscript supplement.
PIHF adapts the evaluate-and-improve structure of policy iteration to a setting where the quality of diagnostic reasoning cannot be reduced to a single programmable reward.
A reasoning-capable model applies the current diagnostic workflow and records the evidence path it followed.
An expert-designed credit system identifies premature closure, weak evidence use, or missed corrective search.
Experts judge whether a proposed revision is clinically sound and whether full-panel performance remains stable.
Accepted feedback becomes a durable rule in a readable artifact, ready for the next model and the next case.
What liteOdyssey executes
The workflow is structured but not linear. New evidence can send the agent back to reframe the pattern, widen the candidates, or challenge its leading diagnosis.
The matched parametric baseline uses the same model and clinical features without the liteOdyssey policy or biomedical tools. Held-out analyses test whether the gain survives beyond policy-development cases.
Disease Recall@1 across 1,243 cases and 722 rare diseases. Matched baseline: 26.5%.
Recall@1 versus 35.1% for the matched baseline.
Recall@1 versus 22.9% for the matched baseline.
The 1,193 held-out cases preserve nearly identical gains, arguing against case-specific policy adaptation.
Recall@1 improved from 16.7% to 20.4%; blinded physician review independently corroborated the differential-diagnosis gain.
The same natural-language policy improved closed- and open-weight reasoning models without changing their weights. The durable artifact is the procedure, not one model release.
This recording shows the reviewer workspace running a full diagnostic walkthrough from a single prompt — phased reasoning, live tool use, and the final structured report. The live reviewer workspace below is interactive.
Reviewer experience
Enter the protected workspace to question the system, inspect the reasoning process, try a phenotype profile, and stress-test where the policy holds or fails.