In short
Stanford researchers have documented a phenomenon they call the mirage effect: state-of-the-art multimodal models generate confident visual descriptions in the complete absence of any image, without flagging the inconsistency. A text-only model with 3 billion parameters outperforms all large frontier models — exceeding 100 billion parameters — on a chest radiology benchmark, without ever having analyzed a single image during evaluation. The B-Clean framework, developed by the same team, reveals that 74 to 77% of questions in three major multimodal benchmarks are solvable without any recourse to vision. Published rankings do not measure visual understanding. They measure something else.
In short: a desert traveler who sees water in the distance is not making it up — their eye captures a coherent, structured, plausible image. But that image results from light refracting through layers of hot air, not from a water surface. Multimodal models do the same: they produce coherent descriptions of images they have never seen. The result is so plausible that benchmarks confuse it with real visual understanding.
The visual mirage — when the model fabricates what it cannot see
Imagine asking a radiologist to comment on a chest X-ray. They respond with confidence: “We observe a right para-hilar opacity consistent with early STEMI.” You inform them there was no X-ray. They do not back down.
This scenario, which would be absurd for a human clinician, precisely describes the behavior documented by Asadi et al. in their MIRAGE paper published on arXiv in March 2026. The study tests multimodal models — VLMs, or vision-language models, capable of jointly processing text and images — by submitting visual questions without providing any image. The results are unambiguous: the most powerful models of the moment generate confident descriptive responses in more than 60% of cases on average, across all question types. Under standard multimodal evaluation protocols, this rate rises to 90–100% [Asadi et al., arXiv:2603.21687v2].
A fundamental distinction: mirage vs. hallucination
The LLM literature has already well documented hallucination: a model invents details in a valid context — it cites a non-existent paper, attributes a wrong quote to a real author. The mirage is a distinct and deeper phenomenon. Where hallucination operates within a coherent epistemic frame (there is a question, a context, an input modality), the mirage constructs an entirely fictional frame. The model does not merely add a false detail to a real image: it constructs a mental image from nothing and treats it as if it were present.
The examples documented by Asadi et al. illustrate the scale of the phenomenon. Models fabricated license plate numbers for unphoto graphed vehicles. Others generated expiration dates on medications absent from any frame. On medical-style prompts, models produced detailed radiological findings — listing anatomical structures, densities, and anomaly profiles — without having accessed any image [Asadi et al., arXiv:2603.21687v2].
The bias toward severity: STEMI, melanoma, carcinoma
On medical benchmarks, susceptibility to non-visual inference reaches 60 to 99% depending on the benchmark tested [Asadi et al., arXiv:2603.21687v2]. This figure would already be alarming if distributed randomly across benign and severe pathologies. It is not. The diagnoses spontaneously generated by models — in the absence of any image — concentrate on serious, urgent-onset pathologies: STEMI (acute coronary syndrome with ST elevation), melanoma, carcinoma. This bias is not documented as intentional. It likely reflects the distribution of medical training data: severe cases are overrepresented in clinical literature, medical Q&A databases, and specialist forums. The model learns that medical visual questions call for severe responses.
Why speak of a mirage rather than an error?
The term “mirage” is not arbitrary. An optical mirage is a coherent, plausible image that appears real but results from a physical mechanism different from the one we believe. A traveler sees water in the desert: the image is there, clear, structured — but it results from the refraction of light through layers of hot air, not from a water surface. A model generating a detailed radiological report produces something structurally correct, lexically appropriate, clinically coherent — but it is not looking at anything. It generates from learned statistical distributions, without any grounding in a real visual signal.
This metaphor has implications for evaluation. If the model produces correct responses on a benchmark while operating in mirage mode, the benchmark score is technically accurate — but it measures something fundamentally different from visual understanding.
The decisive test: removing the image
The method employed by Asadi et al. has a radical elegance: they remove the image. Rather than attempting to detect contamination by analyzing training corpora — a complex undertaking, often impossible with closed models — they directly test whether models can answer correctly without the visual modality that is supposed to be central.
Major benchmarks put to the void
The benchmarks tested cover standard references in the field: VQA-Rad, MicroVQA, MedXpertQA-MM, MMMU-Pro, Video-MMMU, Video-MME. These benchmarks are used by laboratories to compare their multimodal models, by academic publications to position their contributions, and by practitioners to select models suited to their use cases.
On these benchmarks, models tested in “no image” mode — receiving only the question text — achieve scores that retain 70 to 80% of the performance measured with images, sometimes more [Asadi et al., arXiv:2603.21687v2]. This result has a direct reading: a substantial fraction of the score obtained on these benchmarks does not depend on vision. It can be achieved by simple textual inference.
The extreme case: rank 1 without seeing
Perhaps the study’s most striking result is this: on a standard chest X-ray QA benchmark, a model reached rank 1 — first place in the ranking — without accessing any image [Asadi et al., arXiv:2603.21687v2]. The top-ranked model for “X-ray analysis” is a model that never looked at a single X-ray during evaluation.
The 3-billion-parameter model against the giants
The most counter-intuitive result of the study is probably that of the Qwen-2.5 model. This text-only language model — meaning it has no visual processing capability whatsoever — with 3 billion parameters was fine-tuned on ReXVQA, the largest visual question-answering benchmark for chest radiology. The fine-tuning was performed exclusively on the textual question-answer pairs, never including the corresponding images.
Result: this text-only 3-billion-parameter model outperforms all multimodal frontier models tested, including models exceeding 100 billion parameters [Asadi et al., arXiv:2603.21687v2]. It also outperforms human radiologists by more than 10% on average [Asadi et al., arXiv:2603.21687v2].
This result is not a statistical anomaly. It reveals a construction problem: the questions in the ReXVQA benchmark can be answered correctly from text alone. The model does not need to see the images because the answers are accessible via the statistical distributions of textual questions and answers. In other words: the benchmark does not test what it claims to test.
The blind spot of human-machine comparison
This last point deserves attention. When a paper announces that “model X outperforms radiologists on benchmark Y,” the assumption is that the model and radiologists perform the same task — analyzing a medical image. If the benchmark is text-solvable, radiologists are still looking at images. They bring real visual competence to a real visual signal. The model short-circuits the visual signal and exploits textual shortcuts. The comparison does not measure what it appears to measure.
74% of your benchmarks are contaminated
The discovery of the mirage effect led naturally to a question: if text-only models can perform without images, then benchmarks must contain text-solvable questions. In what proportion?
To answer this question, Asadi et al. developed B-Clean — a post-hoc cleanup framework for multimodal benchmarks. The idea is straightforward: for each question in a benchmark, test whether a text-only model can answer it correctly. If so, the question is “compromised” — it does not require vision — and it is removed from the benchmark. The score on the cleaned benchmark (B-Clean) provides a more reliable measure of true visual understanding.
The cleanup numbers
The results are systematic across the three main benchmarks analyzed:
- MicroVQA: 802 questions removed out of 1,042, representing 77.0% of the benchmark eliminated. 240 questions remain after cleanup [Asadi et al., arXiv:2603.21687v2].
- MedXpertQA-MM: 1,486 questions removed out of 2,000, representing 74.3% removed, 514 retained [Asadi et al., arXiv:2603.21687v2].
- MMMU-Pro: 1,302 questions removed out of 1,730, representing 75.3% removed, 428 retained [Asadi et al., arXiv:2603.21687v2].
Three-quarters of questions in three of the most cited benchmarks in the field do not require vision to be solved. This is not a marginal rate, a statistical artifact, or a problem specific to a poorly constructed benchmark. It is a structural and reproducible pattern.
Rankings reconfigure
The impact on model rankings is concrete. On MMMU-Pro after B-Clean cleanup, GPT-5.1 drops from 76.0% to 67.1% accuracy, a fall of 8.9 points. Gemini-3-Pro drops from 81.0% to 72.8%, a fall of 8.2 points [Asadi et al., arXiv:2603.21687v2]. On MicroVQA after cleanup, the drops are described as dramatic across all models tested, though the accessible summaries do not specify exact numerical values [Asadi et al., arXiv:2603.21687v2].
These drops do not disqualify the models. They remain performant. But they reveal that their “raw” scores incorporate a substantial component of out-of-scope textual inference. The question for practitioners is not “is this model good” but “good at what.” On raw benchmarks, the answer is biased.
Why it happens
The mirage effect results from three cumulative mechanisms. First, textual shortcuts in training data: VLMs learn the correlation between textual question and textual answer independently of the visual signal. If question-answer pairs are statistically consistent without the image, the model bypasses vision — it is not trained to lie, it is trained to predict tokens.
Second, benchmark contamination: test data ends up in automatically crawled training corpora [Jacovi et al., EMNLP 2023]. For closed models, the corpus is a trade secret. For open models, detecting contamination across trillions of tokens remains a complex undertaking [Oren et al., ICLR 2024].
Third, structural design flaws in the benchmarks themselves. Building questions that genuinely require the visual signal — non-deducible from text, non-memorizable, visually meaningful — is rarely achieved. Human annotators naturally create exploitable text-response correlations. This is a structural artifact, not negligence.
The Mirage Score introduced by Asadi et al. — (accuracy without image / accuracy with image) × 100% — measures this visual dependence. A score of 90% means the model achieves 90% of its performance without any image. This metric did not exist before MIRAGE: it was assumed that multimodal models used their visual inputs.
Implications for practitioners
The findings of MIRAGE are not academic curiosities. They have direct implications for anyone deploying or evaluating multimodal models in real contexts — medical deployment, model selection, pipeline construction.
Modality ablation as a mandatory test
Before deploying or selecting a multimodal model, a modality ablation must be performed — testing the model without the image — and the Mirage Score calculated. This operation is simple (a few hours on an existing benchmark) but absent from standard evaluation protocols.
In short: the solution is radically simple — remove the image and test whether the model still answers. If yes, the benchmark measures textual reasoning, not vision. This ablation takes a few hours, requires no model modification, and reveals structural biases that public rankings cannot see. Any team selecting a VLM for production should run it before any deployment.
Published rankings (MMMU-Pro, MedXpertQA, MicroVQA) incorporate ~75% non-visual textual inference. The top-ranked model is not necessarily the best visual model — it is the best on a partially text-solvable benchmark.
The case of medical AI
The medical dimension compounds the picture. Mirage diagnoses concentrate on serious urgent-onset pathologies: STEMI, melanoma, carcinoma. A model used in diagnostic assistance that operates in mirage mode generates systematic false positives on the most alarming pathologies. Clinically, a false positive STEMI triggers an urgent cath lab activation. A false negative melanoma leaves a lesion untreated. Certification of medical devices incorporating VLMs should require demonstration that the model actually uses the visual signal.
Key takeaways
-
The mirage effect is empirically documented on current frontier models: more than 60% of responses to visual questions are generated with confidence in the complete absence of any image, across all categories of benchmarks tested [Asadi et al., arXiv:2603.21687v2].
-
74 to 77% of questions in three major multimodal benchmarks (MicroVQA, MedXpertQA-MM, MMMU-Pro) are solvable without recourse to vision — they are removed by the B-Clean framework [Asadi et al., arXiv:2603.21687v2].
-
A text-only model with 3 billion parameters outperforms all frontier VLMs (>100 billion parameters) on a chest radiology benchmark, operating exclusively on textual question-answer pairs, without ever analyzing an image [Asadi et al., arXiv:2603.21687v2].
-
Published rankings partially measure textual inference, not visual understanding. GPT-5.1 and Gemini-3-Pro scores drop by 8 to 9 points on MMMU-Pro after B-Clean cleanup [Asadi et al., arXiv:2603.21687v2].
-
Modality ablation — testing a VLM without its visual inputs — is a simple diagnostic test that should become systematic before any deployment or selection of a multimodal model.