In brief

Large language models were long limited to text. Then, starting in 2021, one idea changed everything: building a shared mathematical space where an image and its written description end up in the same place. This intuition — carried by CLIP — is the foundation on which virtually all of today’s artificial vision rests. But “seeing” for a model does not mean the same thing as for a human, and the data make this clear.


A shared image-text space

In 2021, OpenAI published CLIP (Contrastive Language-Image Pre-training). The idea is simple to describe: simultaneously train an image encoder and a text encoder on 400 million (image, caption) pairs collected from the web, so that both representations land in the same place in a shared vector space.

Training works by contrast: for an associated (image, text) pair, the model maximizes their similarity; for non-associated pairs, it minimizes it. Result: a photo of a cat and the sentence “a ginger cat on a sofa” end up close in this space, while a photo of a mountain moves away.

What this result demonstrates is that linguistic signal alone is sufficient to build generalizable visual representations. CLIP achieves 76.2% accuracy on ImageNet in zero-shot — without having seen a single labeled ImageNet image during training — rivaling a convolutional network trained with full supervision. This shared space becomes the backbone of almost all vision-language architectures in the years that follow.

In short: CLIP does not recognize images the way a human recognizes a face — it computes distances in an abstract space where coherent (image, text) pairs are neighbors. Performance comes from scale: 400 million pairs scraped from the web, not a single human quality judgment.


How to plug vision into an LLM

Once images can be encoded in a vector space, the question becomes: how to connect it to an existing language model?

Three main strategies have emerged.

MLP projection (popularized by LLaVA in 2023) is the most direct: vectors from the visual encoder are projected via a simple network into the LLM’s token space, which then treats them as if they were words. Simple and effective, but the fusion remains shallow — the language model must infer visual relationships from tokens that are, in some sense, foreign to it.

Interleaved cross-attention (Flamingo’s approach, DeepMind, 2022) is more expressive: cross-attention layers are inserted at regular intervals into the language model, allowing each layer to “consult” the visual representations. The main LLM remains frozen — only the added layers are trained. This is computationally more expensive, but the fusion is deeper.

Image tokenization converts the image into discrete tokens injected directly into the text sequence. The image becomes a sequence of “visual words,” processed exactly like text. This approach allows maximum integration, at the cost of some fine-grained information loss.

GPT-4V (OpenAI, 2023) marked the first large-scale integration into a frontier commercial model. The exact architecture was not published, but demonstrated capabilities show multi-step visual reasoning: reading charts, solving mathematical problems with diagrams, understanding technical schematics.

In short: these three strategies differ by depth of coupling. MLP projection = the LLM reads visual tokens as if guessing a foreign language. Cross-attention = the LLM can consult the image at every reasoning floor. Tokenization = the image becomes full-fledged text. The deeper the coupling, the more complex the training.


Audio: a second modality, a different approach

The same logic applies to audio, with some important differences.

Whisper (OpenAI, 2022) demonstrates that a Transformer encoder-decoder trained on 680,000 hours of web audio produces robust speech recognition across dozens of languages. But Whisper remains an audio-to-text converter: it does not fuse modalities at the representation level.

AudioPaLM (Google, 2023) goes further. It merges PaLM-2 (a large language model) and AudioLM (a generative audio model) into a unified architecture: audio is tokenized into discrete tokens that enter the same sequence as text tokens. The model can simultaneously process and generate both text and speech, preserving prosody and speaker identity.

SALMONN (2023, ICLR 2024) takes a complementary approach with two parallel encoders: Whisper for speech, BEATs for non-verbal sounds. Their outputs are fused before injection into an LLM. This enables simultaneous understanding of what is said, the emotions in the voice, and ambient sounds.


Image generation: when the LLM orchestrates

Multimodality is not limited to understanding — it also encompasses generation.

DALL-E 2 (OpenAI, 2022) exploits the CLIP space as a bridge: a text description is encoded in this space, a diffusion model generates the corresponding image embedding, then a decoder synthesizes the final image. Stable Diffusion (2022) makes this approach accessible on consumer hardware via diffusion in a compressed latent space.

The decisive evolution came with DALL-E 3 (2023): ChatGPT automatically reformulates user descriptions into detailed prompts before generation. The LLM becomes a semantic orchestrator — no longer an isolated generation model, but a system where language drives imagery.

GPT-4o (2024) goes even further by integrating image generation directly into the decoder, signaling the convergence of understanding and generation within a single architecture.


Omni-modal models: the next step

The defining structural development of 2024–2025 is the emergence of so-called “omni-modal” models — capable of processing and generating any combination of modalities (text, image, audio, video) within a unified architecture. GPT-4o, Gemini 1.5, and their open-source equivalents embody this trend.

The typical architecture comprises four layers: specialized encoders per modality, projection into a common representation space, cross-modal processing by transformers, and specialized decoders or unified tokenization for generation.

The competition between three approaches remains open: continuous encoding (floating-point vectors), discrete encoding (quantized tokens), and hybrid approaches. Each strategy implies different trade-offs between precision, compute cost, and integration ease.


What these models still don’t really do

There is a notable gap between advertised performance and the reality of visual understanding.

The visual hallucination phenomenon is well documented: vision-language models generate linguistically plausible but factually incorrect image descriptions. Analyses show that as generated sequences lengthen, the influence of the visual signal diminishes in favor of textual statistical regularities — the model “guesses” the continuation rather than “reading” the image.

The MMMU benchmark (11,500 university-level questions across six disciplines) reveals these limits: GPT-4V caps at 56%, Gemini Ultra at 59%. The hardened version of the benchmark (MMMU-Pro, ACL 2025) shows performance drops of 17 to 27 points when questions are embedded in images rather than presented as text — evidence that models exploit the text modality even when the answer is visually obvious.

Another structural problem is the “modality gap”: in CLIP’s vector space, visual and textual representations remain geometrically separated — an artifact of the contrastive loss itself, not of the modalities as such. This gap induces biases: the model exploits linguistic shortcuts rather than information genuinely extracted from the image.

Finally, the tension between unified and modular architectures remains unresolved. Omni-modal models suffer from inter-modality interference during fine-tuning. Mixture-of-experts architectures attempt a middle path, but at the cost of increased training complexity.

In short: the performance posted on benchmarks (56-59%) is not comparable to human vision. On MMMU-Pro, switching the question from text to image drops the score by 17 to 27 points. It is the sign that the model relies mostly on textual cues — when information is purely visual, the system is far less competent than one believes.

Which model to choose by need

ContextRecommendationWhy
Document reading (reports, invoices, scanned PDFs)GPT-4V, Claude 3 OpusOCR robustness + layout understanding, controlled hallucinations on structured text.
Fine-grained image description (captions, alt-text, accessibility)Gemini 1.5 Pro, GPT-4oGood semantic coverage, handles contextual ambiguities.
Audio understanding (transcription + content analysis)Whisper + LLM, or GPT-4oWhisper remains the pure-transcription champion; GPT-4o for joint speech/meaning analysis.
Instruction-driven image generationDALL-E 3, Midjourney v6LLM-orchestrated pipeline, fidelity to complex prompts.
Research/POC on a low budgetLLaVA, Qwen-VL open-sourceNear-zero cost, acceptable quality for fast iteration.

Key takeaways

  • Multimodality rests on a shared image-text mathematical space, made possible by CLIP (2021): images and text descriptions can be directly compared within it.
  • Three fusion strategies dominate: MLP projection (simple, shallow), interleaved cross-attention (expressive, expensive), image tokenization (deep fusion, information loss).
  • Audio integrates via the same logic: sound tokenization and injection into the text sequence, as demonstrated by AudioPaLM and SALMONN.
  • Omni-modal models (GPT-4o, Gemini 1.5) unify all modalities in a single architecture, but competition between encoding approaches remains open.
  • These models present documented limitations: visual hallucinations, dependence on textual regularities, and a structural “modality gap.” The most recent benchmarks show that deep visual understanding remains an unsolved problem.