In Brief
An LLM agent may call the right tool on the first pass and the wrong one on the second, without infrastructure metrics flagging it. LLM observability addresses this problem: tracing each step of the reasoning, measuring output quality, detecting drift. Several specialized platforms have established themselves since 2023 — LangSmith, Arize Phoenix, Helicone, OpenLLMetry, Braintrust — with very different architectures and business models. A standard is emerging on the tooling side: OpenTelemetry’s GenAI semantic conventions, still in Development phase as of March 2026.
Why LLM Observability Is Not Classic Monitoring
In plain terms : monitoring an LLM agent the way you monitor a web server is like checking the quality of a salad by measuring only the fridge temperature. The fridge works fine, the salad can still be rotten.
Traditional APM (Application Performance Monitoring — supervision of application performance) tools are designed for deterministic systems: an SQL (Structured Query Language — database query language) query always produces the same result for the same parameters. LLMs (Large Language Models) are stochastic by nature. Two identical calls can produce different outputs; an agent may succeed at a task in one context and fail in a slightly modified context.
This structural difference creates blind spots: a classic APM can report that latency (response time) is normal and that all tokens (text units billed by the LLM API) were consumed, without detecting that a response is off-topic or that a tool call produced an invalid argument. A direct criticism formulated in the literature: “If your ‘LLM observability’ looks indistinguishable from traditional APM — just with tokens instead of SQL queries — you’re monitoring infrastructure, not AI behavior” (Dev.to, 2025).
Relevant metrics for agents in production therefore combine two distinct layers: infrastructure (latency, cost, error rate) and behavior (output relevance, reasoning quality, tool call consistency).
The Five Main Tools
In plain terms : each platform solves a different piece of the puzzle — some plug in with one line of code but see little, others require more integration but see all the reasoning. The choice depends less on the product than on what you want to see.
LangSmith
Developed by LangChain, LangSmith started in 2023 to solve observability problems internal to the LangChain ecosystem, then expanded to other frameworks. It offers granular tracing (detailed recording of each execution step) (each step of agent reasoning, tool calls, P50/P99 — median and 99th percentile — latency), customizable dashboards, alerts via webhooks (automated HTTP callbacks) or PagerDuty, and automated evaluations with LLM-as-a-judge (one model used as a judge to score another model’s outputs) scoring.
Declared adoption: over 25,000 monthly active teams, 15 billion traces processed, over 300 enterprise customers (Klarna, Snowflake, Boston Consulting Group). LangChain achieved unicorn status in October 2025 after a $125M Series B (Contrary Research).
Limitations: the $0.50 per 1,000 traces pricing is criticized for production volumes. The perception of LangChain ecosystem coupling persists despite the declaration of framework neutrality — le reproche de « vendor lock-in » et de support inégal selon les piles techniques revient régulièrement dans les retours d’utilisateurs. [NON VÉRIFIÉ] : la source secondaire citée jusqu’ici pointait un domaine inexistant, le constat est conservé sans attribution.
Arize Phoenix
Phoenix is Arize’s open-source (freely inspectable and modifiable source code) platform, self-hostable, built on OpenTelemetry (OTel — open standard for distributed application traceability) and the OpenInference specification. No feature gates (paywalled functional restrictions): all functionality is available without a commercial license. It covers LLM calls, tool executions, retrieval (fetching documents to feed the LLM context) operations, and complete agent reasoning. Declared compatibility: OpenAI Agents SDK (Software Development Kit), Claude Agent SDK, LangGraph, LlamaIndex, CrewAI, DSPy, AWS Bedrock, and others.
OpenInference is the attribute specification developed by Arize as a complement to OpenTelemetry: it defines LLM-specific fields (llm.input_messages, llm.token_count.prompt, etc.) and span (time segments representing an operation within a trace) types native to AI workflows (LLM, RETRIEVER, RERANKER, EMBEDDING, TOOL, AGENT, GUARDRAIL).
Community adoption: 7,800+ GitHub stars [NOT VERIFIED: figure from a single date, single source]. Cited as more active in terms of commits than Langfuse in a comparative analysis (ZenML, 2025) [NOT VERIFIED: not corroborated by a second source].
Helicone
Helicone (YC W23) adopts a proxy (intermediate server intercepting requests) architecture rather than SDK: integration amounts to changing the API endpoint (target URL of the API) (one line of code). The technical stack relies on Cloudflare Workers, ClickHouse, and Kafka. Declared added latency: 50–80ms [NOT VERIFIED: self-declared figure].
Features: integrated caching (storing responses to avoid recomputing), rate limiting (capping the number of requests), routing with automatic fallbacks to over 100 providers, automatic cost tracking, latency analytics. The platform positions itself as an “AI Gateway” — its routing component was rewritten in Rust. Free tier: 10,000 requests/month. 2 billion LLM interactions processed (self-declared figure).
The main limitation of the proxy approach: it captures API exchanges but cannot see inside agent reasoning (no trace of intermediate steps if they don’t emit distinct HTTP requests).
OpenLLMetry
OpenLLMetry is an open-source instrumentation library created by Traceloop, an Israeli startup. It extends OpenTelemetry with LLM attributes (model name, prompt/completion tokens, temperature, latency, errors) and is compatible with any existing OTel backend — Datadog, Dynatrace, Langfuse, and others.
In 2025, ServiceNow acquired Traceloop for an estimated $60–80M (Calcalist Tech, 2025) [NOT VERIFIED: the figures for the May 2025 seed round ($6.1M) and the acquisition come from different sources; the exact timeline between the two is not confirmed]. The team announced that OpenLLMetry would remain open-source, and that the technology would be integrated into ServiceNow’s AI Control Tower — a centralized agent governance dashboard. This integration into a proprietary product creates a structural tension whose evolution is not yet documented.
Braintrust
Braintrust differentiates itself through native integration of evaluations into the observability workflow. Its internal database (Brainstore) is designed to quickly query millions of traces. The Loop feature automatically generates prompts, scorers, and datasets from production data.
Funding in February 2026: $80M Series B led by ICONIQ Growth, with participation from a16z and Greylock. Valuation: $800M (Axios, 2026). Certifications: SOC 2 Type II, GDPR, HIPAA.
Identifiable limitation: the platform is proprietary SaaS, with no self-hosted option documented in the sources consulted, which can pose data sovereignty constraints.
The Emerging Standard: OpenTelemetry GenAI
In plain terms : imagine every car brand inventing its own engine diagnostics with its own ports and its own codes. OpenTelemetry GenAI is the equivalent of the standardized OBD-II port — one single format, all tools compatible.
OpenTelemetry (OTel) is the industry’s open standard for tracing, metrics (numerical values aggregated over time), and logs (timestamped event journals) — managed by the CNCF (Cloud Native Computing Foundation, Linux Foundation). Since April 2024, a specific working group (SIG GenAI — Special Interest Group for generative AI) has been developing semantic conventions (shared naming conventions for data fields) for generative AI systems: standardized attributes for LLM calls, agents, embeddings (numerical vector representations of text), and retrieval operations.
The principle is simple: rather than each tool inventing its own field names, all platforms export traces in a common format. This allows switching observability tools without modifying instrumentation (adding probes into code to collect data) code.
As of March 2026, these conventions remain in Development status — not yet Stable. Datadog supports them starting from version 1.37, but retains the old format by default. OpenInference (Arize) and the OTel GenAI conventions coexist without a formally announced convergence, representing a medium-term fragmentation risk (OpenTelemetry blog, 2024).
A documented friction point: “many OTel-based LLM instrumentation libraries don’t strictly adhere to evolving conventions, resulting in vendor-specific solutions” (OpenTelemetry blog, 2024). Convergence is a declared goal, not an achieved state.
September 2026 review. Two movements since March. First, the GenAI conventions have left the main repository for a dedicated one, open-telemetry/semantic-conventions-genai — a sign that the domain has enough volume to live on its own, but also that its stabilisation no longer follows the trunk’s calendar. Second, an extension to agentic systems is under discussion: tracing not only the model call, but tasks, actions, agents, teams, artifacts and memory. The call-level foundation, for its part, is stable in practice — gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons, and the content of exchanges in gen_ai.input.messages / gen_ai.output.messages when the application is configured to record it.
In short: what changed in six months is not the standard’s status — it is still not Stable. It is its scope. Instrumenting an LLM call is a solved problem; instrumenting a fleet of agents handing work to one another is not, and that is exactly where the need is going.
Which Tool for Which Context?
| Context | Recommendation | Why |
|---|---|---|
| LangChain/LangGraph stack, team already invested | LangSmith | Native integration, evals in the same place, reduced onboarding. |
| Data sovereignty (healthcare, defense, public sector) | Arize Phoenix (self-hosted) | Open source, deployable on-premise, zero feature gate. |
| Quick start, multiple providers (OpenAI, Anthropic, Mistral…) | Helicone | 1-line integration, SaaS (Software as a Service — online software) or self-host, automatic fallbacks. |
| Existing observability stack (Datadog — proprietary APM platform), no tool lock-in | OpenLLMetry | Emits in standard OTel, integrates into any backend already in place. |
| Business evaluations at the core of the workflow (user-facing product) | Braintrust | Evals and traces in the same database, Loop generates scorers automatically. |
Comparison Table
| Tool | Model | Architecture | Strengths | Limitations |
|---|---|---|---|---|
| LangSmith | Proprietary SaaS | SDK | LangGraph integration, built-in evals, broad adoption | Volume pricing, LangChain lock-in perception |
| Arize Phoenix | Open source | OTel + OpenInference | Self-hostable, framework-agnostic, no feature gates | Poorly documented adoption figures |
| Helicone | Open source + proxy | HTTP Proxy | 1-line integration, AI gateway, 100+ providers | No trace of internal agent reasoning |
| OpenLLMetry | Open source | Native OTel | Compatible with any OTel backend, standard instrumentation | Post-ServiceNow acquisition uncertainty |
| Braintrust | Proprietary SaaS | SDK + Brainstore DB | Evals integrated into workflow, dedicated trace database | SaaS only, cost at high production volumes |
Real Adoption: Two Figures to Put in Perspective
Two 2025 surveys give very different figures. LangChain’s State of AI Agent Engineering indicates that 89% of respondents have implemented some form of observability for their agents (LangChain, 2025). The Grafana Observability Survey 2025 indicates that LLM observability is used “in production, extensively or exclusively” by only 7% of general respondents (Grafana Labs, 2025).
The gap is explained by selection bias: the LangChain survey targets users already engaged with agents and associated tooling. The Grafana range, covering a general audience, likely better reflects real market adoption.
Limitations and Unresolved Areas
In plain terms : three sticking points that even the best tools do not yet resolve — no referee, a quality measurement that is hard to automate, and data fragmented across several platforms.
Three structural problems remain open.
Absence of independent benchmarks. All available performance figures (added latency, trace reliability, coverage) are self-declared by vendors. No independent comparative study under real production conditions was identified in the sources consulted.
Measuring quality, not just infrastructure. Cost and latency metrics are easily capturable. The quality of an agent’s reasoning — did it make the right choice, call the right tool, produce a correct response — requires evaluations (LLM-as-a-judge, human feedback) that are more complex to automate and interpret. Most tools offer these features, but their production effectiveness is not independently documented.
Fragmentation and total cost of ownership. Production data may reside in one tool (LangSmith), evaluations in another (Braintrust), infrastructure monitoring in a third (Datadog). This dispersion slows iteration. OTel is presented as the universal transport layer to reduce this problem — but the conventions remain unstable and real adoption is heterogeneous.
Key Takeaways
- LLM observability differs from classic APM: the non-deterministic outputs of agents require tracing reasoning, not just infrastructure.
- LangSmith (SaaS, broad adoption), Arize Phoenix (open source, native OTel), Helicone (proxy, minimal integration), OpenLLMetry (standard instrumentation), and Braintrust (integrated evals) cover distinct use cases.
- OpenTelemetry GenAI Semantic Conventions is the standard under construction — in Development status as of March 2026, supported by Datadog, Langfuse, Arize and others, but not yet stabilized.
- No independent benchmark compares these tools in real production: available figures are self-declared.
- The fragmentation risk is real: two specifications coexist (OTel GenAI and Arize’s OpenInference) without a formally announced convergence.