→ Back to Home
Observability

New Relic AI Evaluation Empowers Developers to Ensure Production Safety of AI Applications

New Relic has announced the launch of AI Evaluation, a significant new feature within its AI Observability platform. This tool is designed to assess the behavior of AI applications throughout their entire lifecycle, from development to production. Unlike traditional application monitoring, which primarily focuses on deterministic software metrics such as uptime, latency, and error rates, AI Evaluation delves into the probabilistic nature of AI systems. It provides transaction-level analysis, linking AI performance, model behavior, and business impact to distributed traces. This allows engineering teams to pinpoint whether failures originate from prompts, retrieval systems, models themselves, or underlying infrastructure. The significance of this development for practitioners cannot be overstated. As enterprises increasingly deploy generative AI systems, autonomous agents, and RAG (Retrieval Augmented Generation) pipelines into production, the nature of system failures has fundamentally shifted. An AI system can return a '200 OK' response with normal latency, yet still deliver hallucinated content, expose sensitive data, or make biased decisions. Traditional monitoring tools are blind to these 'semantic failures,' creating a critical visibility gap. AI Evaluation directly addresses this by using an asynchronous "LLM-as-a-judge" service to scan live telemetry and score issues like hallucinations and prompt injections, attaching these scores as attributes to application traces. This integrated view allows engineers to correlate technical and AI-specific signals in a single pane of glass, dramatically improving their ability to diagnose and resolve issues that impact the quality and safety of AI outputs. This move by New Relic fits squarely within the broader, well-established trend of specialized observability for AI systems. The industry has recognized that AI-driven applications introduce unique monitoring requirements that traditional observability tools cannot meet. Concepts like "AI observability" have emerged to specifically address the need for continuous monitoring, tracing, and analysis of AI systems in production to understand their behaviors, execution context, and outputs. This includes measuring quality, safety, costs, and semantic correctness, which are distinct from the metrics of deterministic software. Other vendors have also been expanding their AI observability offerings, with a clear focus on agent evaluation, performance, cost optimization, and controls for harmful actions. The increasing adoption of AI in DevOps, with AI-driven operations becoming a default for modern enterprises, further underscores the necessity for such specialized tools. In practice, this means that DevOps and SRE teams working with AI applications should actively explore and adopt AI-specific observability solutions. Relying solely on traditional APM will leave them vulnerable to critical AI-specific failures that can directly impact business outcomes and user trust. Practitioners should prioritize tools that offer deep visibility into the AI transaction lifecycle, from prompt to response, and can correlate AI-specific quality metrics with underlying infrastructure performance. The ability to evaluate AI behavior at a granular level, detect anomalies like hallucinations or data leaks, and integrate these insights into existing incident response workflows will be crucial. This shift demands a proactive approach to understanding the probabilistic nature of AI systems and investing in the right tooling to ensure their reliable and safe operation in production.
#ai observability#ai evaluation#generative ai#llm#production monitoring#new relic
Read original source