Skip to content

Fundamentals

How it works — the science behind the models

30 articles · 4 sous-catégories
Subcategory
Content type
concept Evaluation

Regression testing for AI systems — measuring an agent's drift

An AI system is not deterministic and evolves over time. Without a reference corpus, there is no drift detection. Method, illustrated figures, limits.

evaluationbaselinereproducibilityregressioncorpusreliabilitymethodology
concept Tokens & context

Contextual degradation is measurable — and follows a law

The loss of coherence in LLMs over long conversations is not a vague bug: it follows a measurable mathematical curve. Two competing models shed light on the phenomenon and its solutions.

contextcontext-rotdegradationlong-contextlambda-RLMlost-in-the-middleattentionwindow
concept Reasoning

The Reasoning Ceiling Under Constraints — 55-60% and No Further

On constrained optimization tasks, LLMs plateau at 55-60% constraint satisfaction, regardless of model. Only reinforcement learning breaks this ceiling.

reasoningconstraintsoptimizationGRPOSFTceilingbenchmarkRL
concept Reasoning

Society of Thought — When Collective Reasoning Emerges Without Instruction

Reasoning models like DeepSeek-R1 spontaneously develop internal debate structures under RL training. These conversational patterns — questioning, contradiction, reconciliation — are causally linked to performance gains.

reasoningmulti-agentsreinforcement-learningemergencedebatedeepseekrlchain-of-thought
concept Memory

Graph memory for agents — beyond the vector

A 2026 survey shows that graphs outperform vectors for agent memory on four dimensions. Taxonomy, tools, and state of the art.

memoryknowledge-graphagentsretrievalembeddingsgraphtaxonomytemporal
decouverte Memory

GradMem — compressing context into memory via gradient descent

GradMem proposes writing context into memory tokens via a few gradient descent steps at inference time. A compression approach that outperforms forward-only methods.

memorycontextcompressiongradient-descentinferencekv-cachegradmem
decouverte Architecture & functioning

HyperAgents — AI agents that modify their own improvement method

HyperAgents extends the Darwin Gödel Machine concept by making the meta-level of agent improvement editable. A step toward recursive self-improvement beyond code.

agentsself-improvementmeta-cognitionhyperagentsdarwin-godel-machinearchitecturerecursion
concept Memory

The 3 memories of an AI agent: semantic, episodic, procedural

An AI agent manages three distinct memory types imported from cognitive psychology. The CoALA framework formalizes them for LLMs by adding working memory. Comparison with MemGPT and recent approaches.

memoryagentsCoALAMemGPTepisodicsemanticproceduralreflectionarchitecture
concept Architecture & functioning

The 4 levels of an AI agent (L0–L3)

The L0–L3 taxonomy organizes the progression of AI agents according to their interaction with the environment: from the bare LLM to collaborative multi-agent systems. Comparison with alternative taxonomies from Yu Huang, HuggingFace, and Google DeepMind.

agentstaxonomymulti-agentsorchestrationcontext-engineeringtool-useautonomyarchitecture
concept Reasoning

LLM Reasoning Techniques Decoded — CoT, ToT, ReAct Explained

An overview of the four major techniques that enable LLMs to reason rather than simply generate: Chain-of-Thought, Tree-of-Thought, ReAct, and the Scaling Inference Law. When to use each.

chain-of-thoughttree-of-thoughtreactreasoninginferencetest-time computepromptingagents
concept Architecture & functioning

LLM Agent Taxonomy — from chatbot to autonomous team

How to classify LLM systems by their actual degree of autonomy: from L0, the tool-less model, to L3, the team of specialized agents. Distinguishing criteria, the perception-planning-action loop, and limits of the classification.

agentstaxonomyautonomymulti-agentscontext-engineeringclassificationperceptionplanning
concept Training

Alignment and RLHF — how to correct an LLM

RLHF is the dominant method for making LLMs conform to human values: a three-phase pipeline, recent alternatives (DPO, KTO), theoretical limits, and the jailbreaking problem by construction.

alignmentrlhfjailbreaksafetyconstitutional-aisycophancy
concept Architecture & functioning

Transformer architecture — the mechanism that changed everything

How attention works, why every major language model relies on the same architecture, and what researchers are still debating today.

transformerattentionself-attentionscalingarchitecture
concept Training

Synthetic data — how LLMs learn from their own outputs

LLMs increasingly generate their own training data. This technique multiplies the capabilities of small models — but it carries a structural risk: the progressive collapse of diversity.

synthetic-datatrainingfine-tuningmodel-collapsescaling-laws
concept Architecture & functioning

Embeddings and semantic search — how LLMs understand meaning

An embedding turns text into coordinates in a mathematical space where proximity expresses kinship of meaning. Understanding this mechanism means understanding why RAG works — and why it fails.

embeddingsvector-storesemantic-searchhnswfaiss
concept Evaluation

LLM Benchmarks — why scores don't tell the whole story

MMLU, GSM8K, Chatbot Arena: how is language model performance actually measured? An overview of evaluation methods, their limitations, and the debates running through current research.

benchmarksevaluationmmluchatbot-arenallm-as-judge
concept Tokens & context

The context window — what it is, why it matters

The context window determines how much information an LLM can process at once. Here is how it works and why it is critical.

context-windowtokenslong-contextattentiontransformer
concept Architecture & functioning

AI image generation — from noise to image

How diffusion models build images from random noise, which players dominate the sector, and why copyright questions remain unresolved.

diffusionimage-generationdall-emidjourneystable-diffusion
concept Architecture & functioning

History of LLMs — five breakthroughs that changed everything

Large language models did not appear overnight. A look back at the five scientific turning points, from 2013 to the present day, that made ChatGPT and its successors possible.

transformerrlhfscaling-lawshistorygptbert
concept Evaluation

Mechanistic interpretability — understanding what happens inside an LLM

How researchers dissect language models to understand their internal mechanisms, and what this reveals about AI safety.

interpretabilitymechanistic-interpretabilitysparse-autoencodercircuitsai-safety
concept Tokens & context

Context window — what LLMs actually see

Models advertise windows of 128,000 or one million tokens. What those numbers hide, how engineers worked around the physical limits, and why information in the middle of a long document is often lost.

contextattentionropeflash-attentionmemorytransformer
concept Architecture & functioning

Mixture of Experts — how large models activate fewer parameters than they contain

MoE models decouple a network's total size from its actual compute cost: only a fraction of parameters is activated for each token processed. Understanding this mechanism, its advantages, and its real limitations.

moesparsescalingefficiencyarchitectureroutingexperts
concept Architecture & functioning

Multimodality — how LLMs learned to see and hear

From CLIP to GPT-4o, a look at the mechanisms that enable large models to process images, audio, and text within a single architecture — and their real limitations.

multimodalvisionaudiogeneration-imagestransformer
concept Architecture & functioning

Chain-of-Thought — how to make an LLM think out loud

Since 2022, a simple prompting technique — asking a model to spell out its steps — has transformed LLM performance on complex tasks. An overview of the mechanisms, extensions, and documented limits.

reasoningchain-of-thoughtcoto1prompting
concept Training

Scaling laws — why bigger means better

Scaling laws describe how LLM performance scales with size, data and compute. From Kaplan to Chinchilla, then to inference-time reasoning: a simple idea with real limits.

scaling-lawschinchillatrainingcomputeparameters
concept Architecture & functioning

Temperature and sampling — controlling an LLM's creativity

Temperature and sampling methods determine whether an LLM produces predictable or surprising responses. Here is how it works.

temperaturesamplingtop-ktop-pnucleus-samplinggeneration
concept Tokens & context

Advanced tokenization — beyond BPE

BPE is not the only tokenization algorithm. WordPiece, Unigram LM, and SentencePiece follow different logics — and the choice of tokenizer has concrete consequences for fairness across languages.

tokenizationbpewordpiecesentencepiecemultilingual
concept Tokens & context

Tokens and tokenization — how LLMs read text

An LLM does not read words but tokens. Understanding tokenization means understanding how these models process language.

tokenstokenizationbpevocabulary
concept Training

LLM training pipeline — from raw text to a useful model

Training a large language model unfolds in three distinct stages: pretraining, supervised fine-tuning, and preference alignment. Each carries its own cost, logic, and limitations.

trainingpretrainingfine-tuningrlhfpipeline
concept Training

LLM alignment — why correcting AI is not enough

RLHF is the dominant method for making LLMs conform to human values. But recent research shows it amplifies certain flaws, marginalizes minority preferences, and cannot eliminate jailbreaking by design.

alignmentRLHFjailbreaksafetyConstitutional AIsycophancy