In Brief

An LLM agent may call the right tool on the first pass and the wrong one on the second, without infrastructure metrics flagging it. LLM observability addresses this problem: tracing each step of the reasoning, measuring output quality, detecting drift. Several specialized platforms have established themselves since 2023 — LangSmith, Arize Phoenix, Helicone, OpenLLMetry, Braintrust — with very different architectures and business models. A standard is emerging on the tooling side: OpenTelemetry’s GenAI semantic conventions, still in Development phase as of March 2026.


Why LLM Observability Is Not Classic Monitoring

In plain terms : monitoring an LLM agent the way you monitor a web server is like checking the quality of a salad by measuring only the fridge temperature. The fridge works fine, the salad can still be rotten.

Traditional APM (Application Performance Monitoring — supervision of application performance) tools are designed for deterministic systems: an SQL (Structured Query Language — database query language) query always produces the same result for the same parameters. LLMs (Large Language Models) are stochastic by nature. Two identical calls can produce different outputs; an agent may succeed at a task in one context and fail in a slightly modified context.

This structural difference creates blind spots: a classic APM can report that latency (response time) is normal and that all tokens (text units billed by the LLM API) were consumed, without detecting that a response is off-topic or that a tool call produced an invalid argument. A direct criticism formulated in the literature: “If your ‘LLM observability’ looks indistinguishable from traditional APM — just with tokens instead of SQL queries — you’re monitoring infrastructure, not AI behavior” (Dev.to, 2025).

Relevant metrics for agents in production therefore combine two distinct layers: infrastructure (latency, cost, error rate) and behavior (output relevance, reasoning quality, tool call consistency).


The Five Main Tools

In plain terms : each platform solves a different piece of the puzzle — some plug in with one line of code but see little, others require more integration but see all the reasoning. The choice depends less on the product than on what you want to see.

LangSmith

Developed by LangChain, LangSmith started in 2023 to solve observability problems internal to the LangChain ecosystem, then expanded to other frameworks. It offers granular tracing (detailed recording of each execution step) (each step of agent reasoning, tool calls, P50/P99 — median and 99th percentile — latency), customizable dashboards, alerts via webhooks (automated HTTP callbacks) or PagerDuty, and automated evaluations with LLM-as-a-judge (one model used as a judge to score another model’s outputs) scoring.

Declared adoption: over 25,000 monthly active teams, 15 billion traces processed, over 300 enterprise customers (Klarna, Snowflake, Boston Consulting Group). LangChain achieved unicorn status in October 2025 after a $125M Series B (Contrary Research).

Limitations: the $0.50 per 1,000 traces pricing is criticized for production volumes. The perception of LangChain ecosystem coupling persists despite the declaration of framework neutrality — le reproche de « vendor lock-in » et de support inégal selon les piles techniques revient régulièrement dans les retours d’utilisateurs. [NON VÉRIFIÉ] : la source secondaire citée jusqu’ici pointait un domaine inexistant, le constat est conservé sans attribution.

Arize Phoenix

Phoenix is Arize’s open-source (freely inspectable and modifiable source code) platform, self-hostable, built on OpenTelemetry (OTel — open standard for distributed application traceability) and the OpenInference specification. No feature gates (paywalled functional restrictions): all functionality is available without a commercial license. It covers LLM calls, tool executions, retrieval (fetching documents to feed the LLM context) operations, and complete agent reasoning. Declared compatibility: OpenAI Agents SDK (Software Development Kit), Claude Agent SDK, LangGraph, LlamaIndex, CrewAI, DSPy, AWS Bedrock, and others.

OpenInference is the attribute specification developed by Arize as a complement to OpenTelemetry: it defines LLM-specific fields (llm.input_messages, llm.token_count.prompt, etc.) and span (time segments representing an operation within a trace) types native to AI workflows (LLM, RETRIEVER, RERANKER, EMBEDDING, TOOL, AGENT, GUARDRAIL).

Community adoption: 7,800+ GitHub stars [NOT VERIFIED: figure from a single date, single source]. Cited as more active in terms of commits than Langfuse in a comparative analysis (ZenML, 2025) [NOT VERIFIED: not corroborated by a second source].

Helicone

Helicone (YC W23) adopts a proxy (intermediate server intercepting requests) architecture rather than SDK: integration amounts to changing the API endpoint (target URL of the API) (one line of code). The technical stack relies on Cloudflare Workers, ClickHouse, and Kafka. Declared added latency: 50–80ms [NOT VERIFIED: self-declared figure].

Features: integrated caching (storing responses to avoid recomputing), rate limiting (capping the number of requests), routing with automatic fallbacks to over 100 providers, automatic cost tracking, latency analytics. The platform positions itself as an “AI Gateway” — its routing component was rewritten in Rust. Free tier: 10,000 requests/month. 2 billion LLM interactions processed (self-declared figure).

The main limitation of the proxy approach: it captures API exchanges but cannot see inside agent reasoning (no trace of intermediate steps if they don’t emit distinct HTTP requests).

OpenLLMetry

OpenLLMetry is an open-source instrumentation library created by Traceloop, an Israeli startup. It extends OpenTelemetry with LLM attributes (model name, prompt/completion tokens, temperature, latency, errors) and is compatible with any existing OTel backend — Datadog, Dynatrace, Langfuse, and others.

In 2025, ServiceNow acquired Traceloop for an estimated $60–80M (Calcalist Tech, 2025) [NOT VERIFIED: the figures for the May 2025 seed round ($6.1M) and the acquisition come from different sources; the exact timeline between the two is not confirmed]. The team announced that OpenLLMetry would remain open-source, and that the technology would be integrated into ServiceNow’s AI Control Tower — a centralized agent governance dashboard. This integration into a proprietary product creates a structural tension whose evolution is not yet documented.

Braintrust

Braintrust differentiates itself through native integration of evaluations into the observability workflow. Its internal database (Brainstore) is designed to quickly query millions of traces. The Loop feature automatically generates prompts, scorers, and datasets from production data.

Funding in February 2026: $80M Series B led by ICONIQ Growth, with participation from a16z and Greylock. Valuation: $800M (Axios, 2026). Certifications: SOC 2 Type II, GDPR, HIPAA.

Identifiable limitation: the platform is proprietary SaaS, with no self-hosted option documented in the sources consulted, which can pose data sovereignty constraints.


The Emerging Standard: OpenTelemetry GenAI

In plain terms : imagine every car brand inventing its own engine diagnostics with its own ports and its own codes. OpenTelemetry GenAI is the equivalent of the standardized OBD-II port — one single format, all tools compatible.

OpenTelemetry (OTel) is the industry’s open standard for tracing, metrics (numerical values aggregated over time), and logs (timestamped event journals) — managed by the CNCF (Cloud Native Computing Foundation, Linux Foundation). Since April 2024, a specific working group (SIG GenAI — Special Interest Group for generative AI) has been developing semantic conventions (shared naming conventions for data fields) for generative AI systems: standardized attributes for LLM calls, agents, embeddings (numerical vector representations of text), and retrieval operations.

The principle is simple: rather than each tool inventing its own field names, all platforms export traces in a common format. This allows switching observability tools without modifying instrumentation (adding probes into code to collect data) code.

As of March 2026, these conventions remain in Development status — not yet Stable. Datadog supports them starting from version 1.37, but retains the old format by default. OpenInference (Arize) and the OTel GenAI conventions coexist without a formally announced convergence, representing a medium-term fragmentation risk (OpenTelemetry blog, 2024).

A documented friction point: “many OTel-based LLM instrumentation libraries don’t strictly adhere to evolving conventions, resulting in vendor-specific solutions” (OpenTelemetry blog, 2024). Convergence is a declared goal, not an achieved state.

September 2026 review. Two movements since March. First, the GenAI conventions have left the main repository for a dedicated one, open-telemetry/semantic-conventions-genai — a sign that the domain has enough volume to live on its own, but also that its stabilisation no longer follows the trunk’s calendar. Second, an extension to agentic systems is under discussion: tracing not only the model call, but tasks, actions, agents, teams, artifacts and memory. The call-level foundation, for its part, is stable in practice — gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons, and the content of exchanges in gen_ai.input.messages / gen_ai.output.messages when the application is configured to record it.

In short: what changed in six months is not the standard’s status — it is still not Stable. It is its scope. Instrumenting an LLM call is a solved problem; instrumenting a fleet of agents handing work to one another is not, and that is exactly where the need is going.


Which Tool for Which Context?

ContextRecommendationWhy
LangChain/LangGraph stack, team already investedLangSmithNative integration, evals in the same place, reduced onboarding.
Data sovereignty (healthcare, defense, public sector)Arize Phoenix (self-hosted)Open source, deployable on-premise, zero feature gate.
Quick start, multiple providers (OpenAI, Anthropic, Mistral…)Helicone1-line integration, SaaS (Software as a Service — online software) or self-host, automatic fallbacks.
Existing observability stack (Datadog — proprietary APM platform), no tool lock-inOpenLLMetryEmits in standard OTel, integrates into any backend already in place.
Business evaluations at the core of the workflow (user-facing product)BraintrustEvals and traces in the same database, Loop generates scorers automatically.

Comparison Table

ToolModelArchitectureStrengthsLimitations
LangSmithProprietary SaaSSDKLangGraph integration, built-in evals, broad adoptionVolume pricing, LangChain lock-in perception
Arize PhoenixOpen sourceOTel + OpenInferenceSelf-hostable, framework-agnostic, no feature gatesPoorly documented adoption figures
HeliconeOpen source + proxyHTTP Proxy1-line integration, AI gateway, 100+ providersNo trace of internal agent reasoning
OpenLLMetryOpen sourceNative OTelCompatible with any OTel backend, standard instrumentationPost-ServiceNow acquisition uncertainty
BraintrustProprietary SaaSSDK + Brainstore DBEvals integrated into workflow, dedicated trace databaseSaaS only, cost at high production volumes

Real Adoption: Two Figures to Put in Perspective

Two 2025 surveys give very different figures. LangChain’s State of AI Agent Engineering indicates that 89% of respondents have implemented some form of observability for their agents (LangChain, 2025). The Grafana Observability Survey 2025 indicates that LLM observability is used “in production, extensively or exclusively” by only 7% of general respondents (Grafana Labs, 2025).

The gap is explained by selection bias: the LangChain survey targets users already engaged with agents and associated tooling. The Grafana range, covering a general audience, likely better reflects real market adoption.


Limitations and Unresolved Areas

In plain terms : three sticking points that even the best tools do not yet resolve — no referee, a quality measurement that is hard to automate, and data fragmented across several platforms.

Three structural problems remain open.

Absence of independent benchmarks. All available performance figures (added latency, trace reliability, coverage) are self-declared by vendors. No independent comparative study under real production conditions was identified in the sources consulted.

Measuring quality, not just infrastructure. Cost and latency metrics are easily capturable. The quality of an agent’s reasoning — did it make the right choice, call the right tool, produce a correct response — requires evaluations (LLM-as-a-judge, human feedback) that are more complex to automate and interpret. Most tools offer these features, but their production effectiveness is not independently documented.

Fragmentation and total cost of ownership. Production data may reside in one tool (LangSmith), evaluations in another (Braintrust), infrastructure monitoring in a third (Datadog). This dispersion slows iteration. OTel is presented as the universal transport layer to reduce this problem — but the conventions remain unstable and real adoption is heterogeneous.


Key Takeaways

  • LLM observability differs from classic APM: the non-deterministic outputs of agents require tracing reasoning, not just infrastructure.
  • LangSmith (SaaS, broad adoption), Arize Phoenix (open source, native OTel), Helicone (proxy, minimal integration), OpenLLMetry (standard instrumentation), and Braintrust (integrated evals) cover distinct use cases.
  • OpenTelemetry GenAI Semantic Conventions is the standard under construction — in Development status as of March 2026, supported by Datadog, Langfuse, Arize and others, but not yet stabilized.
  • No independent benchmark compares these tools in real production: available figures are self-declared.
  • The fragmentation risk is real: two specifications coexist (OTel GenAI and Arize’s OpenInference) without a formally announced convergence.