Practical LLM Evaluation for Production Systems : Measure, monitor, and improve AI system reliability across training and inference

86.13 SGD
会員価格
77.52
English

Singapore Main Store

Available

F05-01 Floor (, )

Product Description

Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applications Free with your book: DRM-free PDF version + access to Packt's next-gen Reader* Key Features Design evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systems Measure quality, safety, grounding, robustness, and production readiness with practical metrics Apply unified evaluation methods to text, multimodal, and agentic AI systems Book DescriptionModern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks. Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives. Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production. By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence. *Email sign-up and proof of purchase required What you will learn Design repeatable evaluation pipelines for LLM system

Available to Order

Usually dispatches within 3-5 business days

While every attempt has been made to ensure stock availability, occasionally we may run out of stock at our stores.

ご注文金額 50.00 SGD以上で国内送料無料

Discount is applied at checkout.

Recently Viewed Items

Related Products