Observability for Large Language Models: Site Reliability and Chaos Engineering for AI at Scale - Ankush Sharma - cover

LIBRO INGLESE

Observability for Large Language Models: Site Reliability and Chaos Engineering for AI at Scale

di Ankush Sharma (Autore)

APress, 2026

Recensioni: 0/5

54,99 € +550 punti Effe

Venditore: Feltrinelli

Articolo acquistabile con Carta del Docente

Articolo acquistabile con Carta Cultura Giovani e Carta del Merito

Articolo acquistabile con Carta della Cultura

Paga con Klarna in 3 rate senza interessi per ordini superiori a 39 €

Dati e Statistiche

Salvato in 0 liste dei desideri

Observability for Large Language Models: Site Reliability and Chaos Engineering for AI at Scale

Disponibilità in 3 settimane

54,99 €

Disponibilità in 3 settimane

Descrizione

This book is a comprehensive guide designed to equip engineers, data scientists, and AI practitioners with the principles, tools, and strategies needed to ensure reliability, performance, and accountability in Large Language Models (LLMs). The book begins by laying the groundwork with the foundations of observability, introducing LLMs, their significance in modern AI, and the critical role observability plays in maintaining robust systems. It then explores SRE principles, service level objectives, and incident response, while distinguishing the unique observability challenges that arise in AI and ML systems. Building on this foundation, the book dives into measuring performance, from defining SLOs tailored for LLMs to monitoring computational and token-level metrics. Readers gain practical insights into structured logging, debugging, and distributed tracing methods that provide visibility into complex LLM workflows. Scaling challenges are addressed through strategies for cross-model observability, autoscaling, latency reduction, and fault-tolerant infrastructure design. The book further explores chaos engineering, guiding readers through resilience testing in LLMs and the automation of chaos experiments in CI/CD pipelines. Finally, it highlights monitoring, retraining, and ethical considerations in AI observability, including governance, privacy, and accountability. In conclusion, this book provides a holistic roadmap to building reliable, transparent, and future-ready LLM systems. What you will learn: How to design observability pipelines for LLMs, including token-level logging, prompt tracing, and latency analysis. Techniques for applying chaos engineering principles to test LLM robustness under stress and failure scenarios. Methods for building SLOs, SLAs, and dashboards tailored to inference quality and model reliability. Strategies for monitoring hallucinations, drift, bias, and ethical failures in real-time. Who this book is for: This book is for AI infrastructure engineers, SREs, machine learning platform teams, and applied AI practitioners deploying or maintaining LLM-based applications.

Dettagli

Autore:

Ankush Sharma
Editore:

APress
Anno:

2026
Rilegatura:

Paperback / softback
Pagine:

121 p.

Testo in English

Dimensioni:

254 x 178 mm

EAN:

9798868828263

Informazioni e Contatti sulla Sicurezza dei Prodotti

Le schede prodotto sono aggiornate in conformità al Regolamento UE 988/2023. Laddove ci fossero taluni dati non disponibili per ragioni indipendenti da Feltrinelli, vi informiamo che stiamo compiendo ogni ragionevole sforzo per inserirli. Vi invitiamo a controllare periodicamente il sito www.lafeltrinelli.it per eventuali novità e aggiornamenti.
Per le vendite di prodotti da terze parti, ciascun venditore si assume la piena e diretta responsabilità per la commercializzazione del prodotto e per la sua conformità al Regolamento UE 988/2023, nonché alle normative nazionali ed europee vigenti.

Per informazioni sulla sicurezza dei prodotti, contattare productsafety@feltrinelli.it

Observability for Large Language Models: Site Reliability and Chaos Engineering for AI at Scale

Descrizione

Dettagli

Questo prodotto lo trovi anche