Goktug Ozkan
Most medical AI benchmarks ask a simple question: did the model get the right answer? MedFailBench flips that around and asks which safety boundary broke. Instead of scoring correctness, it labels errors by severity on a one-to-five scale and by the type of safety gate that failed, things like missing an urgent escalation, giving unsafe dosing advice remotely, reassuring a patient toward an unsafe discharge, or fabricating evidence.
Worth being clear about the scope here. This early public release holds 44 synthetic, clinician-reviewed cases with a severity rubric, a safety-gate taxonomy, and a live leaderboard preview. It contains no patient data, no clinical validation claims, and no model rankings. It ships under open licenses with a citable DOI.
So this is a small, deliberately cautious first pass at cataloguing how medical AI fails, not a finished evaluation. Read the paper and release notes for the taxonomy and its stated limits.
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains…
Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening
arXiv (cs.CV) · July 16, 2026Weakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based Diffusion
arXiv (cs.CV) · July 16, 2026Video = World + Event Stream
arXiv (cs.AI) · July 16, 2026Man, Machine, and Masterpiece: Artistic Ownership in the AI Era
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.