Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · July 16, 2026

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

Goktug Ozkan

Most medical AI benchmarks ask a simple question: did the model get the right answer? MedFailBench flips that around and asks which safety boundary broke. Instead of scoring correctness, it labels errors by severity on a one-to-five scale and by the type of safety gate that failed, things like missing an urgent escalation, giving unsafe dosing advice remotely, reassuring a patient toward an unsafe discharge, or fabricating evidence.

Worth being clear about the scope here. This early public release holds 44 synthetic, clinician-reviewed cases with a severity rubric, a safety-gate taxonomy, and a live leaderboard preview. It contains no patient data, no clinical validation claims, and no model rankings. It ships under open licenses with a citable DOI.

So this is a small, deliberately cautious first pass at cataloguing how medical AI fails, not a finished evaluation. Read the paper and release notes for the taxonomy and its stated limits.

From the arXiv (cs.AI) abstract

Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains…


More Artificial Intelligence papers