LangWatch
LangWatch is an LLM engineering platform for teams shipping AI agents to production. It combines agent simulation testing, LLM evaluation, observability, prompt management, and AI governance in one place, so you can catch regressions before release and debug what happens in live traffic.
The platform centers on simulation-based testing: realistic multi-turn conversations run against your agent in text or voice, with judge agents scoring outcomes against criteria you define. You can run scenarios locally, in CI/CD, or from the UI, then turn production traces into new test cases when something breaks in the wild.
LangWatch also traces every LLM call, tool invocation, and retrieval step with OpenTelemetry-native instrumentation. Python, TypeScript, and Go SDKs plug into LangGraph, LangChain, CrewAI, OpenAI Agents, and dozens of other frameworks. Deploy on managed cloud, hybrid, VPC, or fully self-hosted with Docker or Kubernetes.
Run multi-turn agent simulations with LLM-powered user simulators and judge agents that score pass or fail
Evaluate offline in CI/CD or online on production traffic with built-in RAGAS, toxicity, and LLM-as-a-judge evaluators
Trace every LLM call, tool use, and retrieval with OpenTelemetry spans, cost tracking, and a visual trace explorer
Version, deploy, and A/B test prompts as code with GitHub sync and deployment-stage tags
Self-host on Docker or Kubernetes, or run managed cloud with SSO, RBAC, SCIM, and ISO 27001 compliance
Free Developer tier with no credit card required to start sending events.
Open-source Scenario SDK for agent simulation testing in Python and TypeScript.
Deploy managed cloud, hybrid, VPC, or fully self-hosted depending on data residency needs.
25+ framework and model provider integrations with OpenTelemetry-native tracing.
Growth plan usage-based pricing adds cost quickly beyond included event limits.
Enterprise features like SSO, SCIM, and custom SLAs require a sales conversation.
Is LangWatch free to use?
Yes. The Developer plan is free forever with 50,000 events per month, 14-day data access, two users, and community support via GitHub and Discord. No credit card is required to sign up.
How much does LangWatch cost for teams?
The Growth plan costs $34 per core seat per month and includes 200,000 events, 30-day data retention, unlimited simulations and evals, and private Slack or Teams support. Usage beyond included limits is billed at $6 per 100,000 events and $4 per GB for extended retention.
Can LangWatch be self-hosted?
Yes. LangWatch runs fully self-hosted with Docker Compose on your own ClickHouse infrastructure, so data stays in your environment. Enterprise self-hosting adds SSO, RBAC, SLAs, and dedicated support.
What frameworks does LangWatch integrate with?
LangWatch offers Python and TypeScript SDKs plus OpenTelemetry support, with first-class integrations for LangGraph, LangChain, CrewAI, OpenAI Agents, Pydantic AI, Vercel AI SDK, AWS Bedrock, Azure OpenAI, and Vertex AI.
Does LangWatch support voice AI agent testing?
Yes. LangWatch can simulate voice agents end to end with providers like ElevenLabs and OpenAI Realtime, including latency metrics and background noise or interruption injection during test runs.

