👉 Access our AI Starter Apps & join hundreds of serious AI builders in our community: https://www.theaiautomators.com/?utm_...
🔗 Langfuse
Website: https://langfuse.com/
GitHub (MIT licensed): https://github.com/langfuse/langfuse
🔗 Coding Agent Plugins
Claude Code: https://langfuse.com/integrations/dev...
Pi Agent: https://langfuse.com/integrations/dev...
Hermes Agent: https://langfuse.com/integrations/oth...
Our Instructions on patching the Hermes Langfuse plugin: https://github.com/theaiautomators/vi...
🔗 Langfuse Agent Skill and MCP server
Agent Skill: https://github.com/langfuse/skills
MCP server: https://langfuse.com/docs/api-and-dat...
🔗 Setup Instructions
Overview: https://langfuse.com/self-hosting
Software with a model in the loop doesn't crash when it's wrong. It just carries on confidently. All of the logs are clean, the tests pass, there are no specific errors thrown, and the answer is still wrong. Meanwhile your costs may have tripled overnight, or response quality has dropped even though you haven't changed the model, and everything on the surface still looks fine.
So what you actually need is X-ray vision into the app. A complete record of every request showing you what the model was handed and when, what it did with that information, what that turn cost, and whether the answer was actually any good or not. That's essentially what Langfuse is. It's open source and MIT licensed, so it either runs on your own machine or free on their hosted tier, and it plugs straight into the coding agent you already work in.
In this video I go through everything you need to know about it. I get it up and running in Claude Code, in the Pi agent and in the Hermes agent, where you very much do not get the same visibility in all three, and then give it the full feature tour connected to the custom AI app we're building on this channel. This video isn't sponsored. We don't do sponsored content on this channel.
⏱️ Timestamps:
0:00 Pi Coding Agent
5:25 Hermes Agent
6:53 Claude Code
8:23 Tracing A Custom AI App
10:04 Why You Need This
13:00 What A Good Trace Looks Like
18:31 Prompt Management
20:10 Scores, Feedback And Evals
23:13 Sensitive Data
23:59 Hosting Options
#AI #AIAgents #Langfuse #LLMObservability #Observability #Evals #LLMEvaluation #Tracing #ClaudeCode #Codex #HermesAgent #PiAgent #AgenticRAG #PromptManagement #OpenSource #AIArchitects #AIBuilder