Skip to content

Issues

Questions that go beyond the technical

12 articles · 3 sous-catégories
Subcategory
Content type
comparatif Evaluation

What the labs publish about their models — and what they leave out

A same-day survey of the model pages at Anthropic, OpenAI and Google: four facts decide an architecture, and none of the three publishes all four.

evaluationdocumentationtransparencyanthropicopenaigooglearchitecturemodels
analyse Economics

The real cost of a multi-agent system — an order of magnitude higher

A solo agent costs $9 per session. A multi-agent system with an orchestrator costs $200. Where does that 22× factor come from, and when is it justified? An anatomy of hidden costs: coordination tokens, context recreation, and verification loops.

multi-agentscosttokensorchestrationprompt-cachingeconomicsproductionscaling
analyse Evaluation

LLM judges are three times more permissive than humans

Quantitative data on the LLM-as-judge problem: false positive rates above 75%, balanced accuracy of 69%, and regression calibration that cuts error by 72%. A numbers-first look at a methodological crisis.

evaluationLLM-as-Judgefalse-positivescalibrationmetrologyAI-Scientistbenchmarkpermissiveness
analyse Evaluation

Mirage — when LLMs "see" without looking

State-of-the-art multimodal models produce confident descriptions of images they do not analyze. 74% of multimodal benchmarks are contaminated. Multimodal evaluation is structurally broken.

multimodalVLMbenchmarkcontaminationevaluationmiragevisionradiology
guide Deployment

The 14 failure modes of AI agents

The MAST taxonomy identifies 14 failure modes in multi-agent LLM systems, grouped into three categories. A practitioner's guide for anticipating and diagnosing agent failures.

agentsfailureMASTmulti-agentdebuggingproductiontaxonomyreliability
analyse Governance

Agentic AI — catalyst for a next intelligence explosion?

Evans, Bratton and Agüera y Arcas (Google/UChicago, 2026) argue that the next intelligence explosion will not be a monolithic mind but a distributed society of human and artificial agents. Analysis of their thesis and its limits.

agentsagentic-aisingularitycollective-intelligencemulti-agentsgovernancealignmentprospective
analyse Governance

Made by AI — so what?

All content on this site is produced with AI assistance. Here's why we own it, what it concretely changes, and why transparency is the only honest position.

transparencygenerative-aimethodethicscontent-productionsourcingverificationauthenticity
concept Evaluation

Why your agent does not know whether it succeeded

Between 27% and 78% of the successes reported by agent benchmarks hide procedural violations. The binary success/failure criterion is structurally insufficient — and the alternatives remain open problems.

evaluationagentsbenchmarksverificationtrajectorymonitoringmulti-agentsquality
concept Evaluation

Why LLMs fail at self-evaluation

Large language models are used to evaluate themselves and other LLMs. Four systemic biases explain why this paradigm breaks down — and why correcting them is harder than it appears.

evaluationbiasLLM-as-Judgecalibrationself-evaluationoverconfidencedebiasingmetrology
concept Safety

LLM Guardrails — safety as architecture

Guardrails are multi-layered control mechanisms that frame the behavior of autonomous agents. They do not limit the agent's capabilities: they make them deployable in real-world conditions.

guardrailssafetyagentsvalidationfilteringleast privilegepydanticcheckpoint
concept Deployment

Human-in-the-Loop — the human and the agent, a necessary pair

The Human-in-the-Loop pattern structures human oversight of agentic systems: when to intervene, how to delegate, and how humans remain essential without blocking automation.

human-in-the-loopsupervisionrlhfescalationautonomyagentsdeploymentalignment
concept Governance

Regulating AI — why Europe, the United States, and China don't speak the same language

Three powers, three regulatory philosophies. Europe classifies risks, China controls content, and the United States oscillates depending on the administration.

governanceregulationeu-ai-actdigital-omnibussafetyethicsgeopolitics