Trace ingestion from observability tools
Connects to Langfuse, Braintrust, Datadog, or a custom trace source and ingests conversation data automatically on a nightly cadence.
Basalt helps teams improve AI agents from real user traces by analyzing failures and generating validated PRs for observability-driven workflows.
Basalt is a product for improving AI agents from real user interactions. It connects to observability or trace sources, reads the previous day’s conversations, and identifies recurring failure patterns that point to where an agent’s behavior needs adjustment.
When it finds a problem, Basalt generates a targeted pull request that can change a prompt, tool definition, or system instruction. Before the PR is opened, the proposed fix is validated against a golden dataset so existing behavior is checked for regressions.
Connects to Langfuse, Braintrust, Datadog, or a custom trace source and ingests conversation data automatically on a nightly cadence.
Uses Claude to cluster conversations, detect failure patterns, and identify which interactions point to quality gaps.
Generates a focused PR that can update a prompt, tool definition, or system instruction based on the detected issue.
Runs each generated fix against a golden dataset before the PR is opened, blocking changes that introduce regressions.
Positions the product as an automated review loop that does not rely on LLM-as-a-judge or manual annotation.
Sends one high-quality PR every morning from the previous day’s user behavior, keeping the workflow tied to recent interactions.
Teams that already record traces in observability tools can connect those sources and let Basalt scan conversations overnight for recurring errors.
When the same misunderstanding or tool-use mistake keeps appearing, Basalt can generate a focused PR that targets the affected behavior pattern.
Teams that maintain golden datasets for expected behavior can use Basalt’s validation step to check that a proposed fix does not regress known good cases.
Organizations that want prompt or tool updates without hand-writing every change can use Basalt to propose modifications based on real user behavior.
Teams looking for a lightweight improvement loop can review one daily PR rather than manually sorting through every interaction.
Basalt ingests conversations from connected observability sources, analyzes them overnight, and turns repeated failure patterns into a targeted PR. The homepage says it can generate prompt updates, tool definition changes, or system instruction changes, then validate the fix against a golden dataset before opening the PR.
The homepage explicitly names Langfuse, Braintrust, Datadog, and a custom trace source as connection options. It presents Basalt as working on traces from those sources rather than requiring manual annotation.
Basalt is positioned around improving agent behavior from real user conversations. The page highlights system prompts, tool definitions, and few-shot examples as the parts it can change when it finds a failure pattern.
The homepage says Basalt runs fixes against a golden dataset before a PR is opened and blocks the PR if regressions appear. It also states there is no LLM-as-a-judge and no manual annotation in the main workflow.
The pricing section says you only pay when you merge a PR, with the first 3 PRs free and no subscription or commitment. It also notes pricing applies per merged PR per repository and that volume discounts are available for teams merging 20+ PRs per month.
blop is a QA agent that writes browser tests as code in your repo, runs them in CI, clusters repeated failures, and can open PRs to fix broken tests.
Orca is an Agent Development Environment for shipping with coding agents, running multiple CLI agents in parallel across isolated worktrees, with desktop and mobile workflows.
Firebase Studio is a web-based workspace for full-stack app development with Gemini-assisted coding, app previews, cloud emulators, collaboration, and browser deployment.
AI Magicx is a unified AI workspace for chat, image, video, voice, music, email and developer tasks, helping teams and creators manage multiple models in one place.
Paper is a design tool that connects canvas, code, and AI agents so teams can create, share, and ship work in one workflow. Includes desktop app and MCP access.
RLAMA is a local AI platform for building RAG systems and intelligent agents on macOS, Linux, and Windows, with HTTP API support.