Goktug Ozkan
Most medical AI benchmarks ask a simple question: did the model get the right answer? MedFailBench flips that around and asks which safety boundary broke. Instead of scoring correctness, it labels errors by severity on a one-to-five scale and by the type of safety gate that failed, things like missing an urgent escalation, giving unsafe dosing advice remotely, reassuring a patient toward an unsafe discharge, or fabricating evidence.
Worth being clear about the scope here. This early public release holds 44 synthetic, clinician-reviewed cases with a severity rubric, a safety-gate taxonomy, and a live leaderboard preview. It contains no patient data, no clinical validation claims, and no model rankings. It ships under open licenses with a citable DOI.
So this is a small, deliberately cautious first pass at cataloguing how medical AI fails, not a finished evaluation. Read the paper and release notes for the taxonomy and its stated limits.
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains…
MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction
arXiv (cs.CV) · July 17, 2026FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
arXiv (cs.CV) · July 17, 2026Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
arXiv (cs.LG) · July 17, 2026PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization