What will I get if I subscribe to this Certificate?

When you enroll in the course, you get access to all of the courses in the Certificate, and you earn a certificate when you complete the work. Your electronic Certificate will be added to your Accomplishments page - from there, you can print your Certificate or add it to your LinkedIn profile.

Building Reliable LLM Systems

Sparen Sie mit 40% Rabatt auf 3 Monate Coursera Plus bei den Fähigkeiten, die Sie zum Strahlen bringen. Jetzt sparen

kurs ist nicht verfügbar in Deutsch (Deutschland)

Wir übersetzen es in weitere Sprachen.

Building Reliable LLM Systems

Dieser Kurs ist Teil von LLM Engineering That Works: Prompting, Tuning, and Retrieval (berufsbezogenes Zertifikat)

Dozent: Professionals from the Industry

Bei enthalten

Mehr erfahren

5 Module

Verschaffen Sie sich einen Einblick in ein Thema und lernen Sie die Grundlagen.

Stufe Mittel

Empfohlene Erfahrung

2 Wochen zu vervollständigen

unter 10 Stunden pro Woche

Flexibler Zeitplan

In Ihrem eigenen Lerntempo lernen

5 Module

Verschaffen Sie sich einen Einblick in ein Thema und lernen Sie die Grundlagen.

Stufe Mittel

Empfohlene Erfahrung

2 Wochen zu vervollständigen

unter 10 Stunden pro Woche

Flexibler Zeitplan

In Ihrem eigenen Lerntempo lernen

Was Sie lernen werden

Build scripts with lexical/semantic metrics to evaluate LLMs, diagnose hallucinations, and balance vector-search recall against latency.
Apply hypothesis testing, confidence intervals, and significance metrics to evaluate model accuracy and validate results from A/B experiments.
Utilize parameterized SQL and data manipulation to segment user logs, calculate retention, and securely retrieve large-scale datasets.
Analyze LLM performance gaps to prioritize technical fixes and implement remediation measures for production-level reliability.

Kompetenzen, die Sie erwerben

Kategorie: Performance Tuning
Kategorie: Debugging
Kategorie: Retrieval-Augmented Generation
Kategorie: Statistical Methods
Kategorie: Statistical Hypothesis Testing
Kategorie: Performance Testing
Kategorie: LLM Application
Kategorie: MLOps (Machine Learning Operations)
Kategorie: Artificial Intelligence and Machine Learning (AI/ML)
Kategorie: SQL
Kategorie: Data-Driven Decision-Making
Kategorie: Statistical Analysis
Kategorie: Model Evaluation
Kategorie: Large Language Modeling

Werkzeuge, die Sie lernen werden

Kategorie: Vector Databases
Kategorie: Python Programming
Kategorie: Query Languages

Wichtige Details

Zertifikat zur Vorlage

Zu Ihrem LinkedIn-Profil hinzufügen

Kürzlich aktualisiert!

März 2026

Bewertungen

14 Zuweisungen¹

KI-bewertet siehe Haftungsausschluss

Unterrichtet in Englisch

Erfahren Sie, wie Mitarbeiter führender Unternehmen gefragte Kompetenzen erwerben.

Weitere Informationen zu Coursera für Unternehmen

Logos von Petrobras, TATA, Danone, Capgemini, P&G und L'Oreal

Erweitern Sie Ihr Fachwissen im Bereich Machine Learning

Dieser Kurs ist Teil der Spezialisierung LLM Engineering That Works: Prompting, Tuning, and Retrieval (berufsbezogenes Zertifikat)

Wenn Sie sich für diesen Kurs anmelden, werden Sie auch für dieses berufsbezogene Zertifikat angemeldet.

Lernen Sie neue Konzepte von Branchenexperten
Gewinnen Sie ein Grundverständnis bestimmter Themen oder Tools
Erwerben Sie berufsrelevante Kompetenzen durch praktische Projekte
Erwerben Sie ein Berufszertifikat von Coursera zur Vorlage

In diesem Kurs gibt es 5 Module

Building Reliable LLM Systems is a comprehensive course for AI practitioners looking to move beyond basic models and create production-grade applications. While getting an LLM to generate text is easy, ensuring a consistently accurate, relevant, and trustworthy output is a significant engineering challenge. This course provides a systematic framework for tackling the entire lifecycle of LLM reliability.

You will start by learning to quantitatively evaluate model performance using a suite of lexical and semantic metrics, such as BLEU, ROUGE-L, and cosine similarity. You’ll dive deep into debugging, using log analysis and data manipulation to uncover the root causes of critical failures, such as hallucinations, by correlating them with retrieval system performance. The course emphasizes statistical rigor, teaching you to design and analyze A/B tests, apply hypothesis testing, and calculate confidence intervals to prove the significance of your optimizations. Finally, you’ll optimize the foundational data layers, learning to tune SQL queries and vector search parameters to achieve the perfect balance between recall and latency.

This module lays the groundwork for quantitative Large Language Mode (LLM) evaluation. Learners will discover why relying on intuition to judge model performance is unsustainable and explore the foundational metrics used to create automated, objective evaluation systems. We will cover both lexical similarity metrics (like BLEU and ROUGE-L) that assess text structure and semantic metrics (like cosine similarity) that capture meaning. By the end of this module, learners will have the conceptual understanding and practical code to build their first automated evaluation script.

Das ist alles enthalten

8 Videos3 Lektüren3 Aufgaben3 Unbewertete Labore

8 VideosInsgesamt 44 Minuten

How to Compute Lexical Metrics: BLEU & ROUGE-L in Python?6 Minuten
How to Compute Semantic Similarity with Embeddings?6 Minuten
Why Guess When You Can Know? The Case of the "Better" Prompt5 Minuten
The Language of Experimentation: Hypotheses, P-Values, and Power5 Minuten
Running the Numbers: A/B Test Analysis in Python7 Minuten
From Report to Action: The Optimization Loop3 Minuten
Case Study: Benchmarking a Sentiment Analyzer6 Minuten
Scripting Your First Evaluation Report6 Minuten

3 LektürenInsgesamt 17 Minuten

A Guide to LLM Evaluation: Lexical and Semantic Metrics5 Minuten
Designing a Fair Race: A/B Testing for LLMs7 Minuten
Building a Reproducible Evaluation Workflow5 Minuten

3 AufgabenInsgesamt 50 Minuten

Knowledge Check: Choosing Your Metrics10 Minuten
Knowledge Check: Statistical Testing Concepts10 Minuten
Build Your LLM Evaluation Toolkit30 Minuten

3 Unbewertete LaboreInsgesamt 128 Minuten

Building Your First Automated Evaluation Script60 Minuten
Statistical Significance Testing60 Minuten
Planning Your Optimization Strategy8 Minuten

When a production chatbot starts giving incorrect answers, how do you find the problem and fix it? This module equips AI practitioners, ML engineers, and data analysts with the essential skills for debugging production LLMs. Go beyond theory and learn the systematic, data-driven workflow that professionals use to solve the critical problem of AI hallucinations. You will be equipped to transition from merely observing AI failures to expertly diagnosing and resolving them.

Das ist alles enthalten

5 Videos3 Lektüren3 Aufgaben2 Unbewertete Labore

5 VideosInsgesamt 29 Minuten

Why Logs Matter: The Air Canada Case?6 Minuten
Calculating Retention in Pandas6 Minuten
Why RAG Fails: The Root of Hallucination?6 Minuten
Correlating Errors with Retrieval in Pandas6 Minuten
Visualizing the Proof in Matplotlib5 Minuten

3 LektürenInsgesamt 28 Minuten

Anatomy of a Log File8 Minuten
The Engineering Brief: From Analysis to Action10 Minuten
Authoring the Engineering Brief10 Minuten

3 AufgabenInsgesamt 40 Minuten

Knowledge Check: Retention Metrics5 Minuten
Knowledge Check: Communicating Findings5 Minuten
LLM Diagnostics Report30 Minuten

2 Unbewertete LaboreInsgesamt 120 Minuten

Lab 1: Segmenting Users & Finding the Drop60 Minuten
Lab 2: Proving the Root Cause60 Minuten

When making high-stakes deployment decisions, a simple accuracy score is not enough. This module equips you with the statistical methods to rigorously validate LLM performance improvements. By the end of this module, you will be able to move beyond subjective "it seems better" evaluations to confidently state, "we can prove it's better," ensuring every deployment decision is backed by sound statistical evidence.

Das ist alles enthalten

5 Videos2 Lektüren3 Aufgaben3 Unbewertete Labore

5 VideosInsgesamt 30 Minuten

Why Single Scores Lie8 Minuten
Calculating Wilson Intervals in Python4 Minuten
Why Gut Feelings Fail in A/B Testing6 Minuten
Running a Chi-Square Test in Python6 Minuten
Visualizing Confidence with Matplotlib5 Minuten

2 LektürenInsgesamt 14 Minuten

Core Concepts: Confidence and Significance8 Minuten
Storytelling with Statistical Visuals6 Minuten

3 AufgabenInsgesamt 40 Minuten

Confidence Intervals Quiz5 Minuten
Communicating Results Quiz5 Minuten
LLM Evaluation Report30 Minuten

3 Unbewertete LaboreInsgesamt 110 Minuten

Lab 1: Quantifying Model Accuracy20 Minuten
Lab 2: Validating a Model Improvement30 Minuten
Lab 3: Create a Comparison Chart60 Minuten

In the world of large-scale AI, slow queries and inefficient search can bring a system to its knees. This module provides the critical skills to prevent that, focusing on practical database and vector search optimization techniques. By the end of this module, you will be equipped to systematically analyze and optimize production retrieval systems, ensuring your AI applications are not only powerful but also fast and reliable.

Das ist alles enthalten

4 Videos3 Lektüren4 Aufgaben3 Unbewertete Labore

4 VideosInsgesamt 26 Minuten

From Inefficient to Optimized7 Minuten
The Recall vs. Latency Trade-Off5 Minuten
Tuning an HNSW Index8 Minuten
Beyond One-Off Tests: The Need for Continuous Benchmarking5 Minuten

3 LektürenInsgesamt 25 Minuten

Secure and Efficient Query Patterns10 Minuten
Understanding Vector Search Parameters10 Minuten
Core Metrics of a Benchmarking Framework5 Minuten

4 AufgabenInsgesamt 85 Minuten

SQL Security and Patterns15 Minuten
Parameter Tuning Scenarios Quiz15 Minuten
Interpreting Benchmark Results10 Minuten
Submit Your Performance Optimization Report45 Minuten

3 Unbewertete LaboreInsgesamt 140 Minuten

Identifying Slowest Queries using Parameterized SQL20 Minuten
Tune HNSW Parameters for Recall and Latency60 Minuten
Create an Automated Benchmarking Suite60 Minuten

In this module, you will conduct an end-to-end performance audit comparing two LLM variants using an A/B test dataset. You will implement a pipeline to calculate key performance metrics, including lexical and semantic similarity, and use statistical A/B testing to validate model improvements. The project culminates in a comprehensive report where you will correlate hallucination rates with retrieval logs and synthesize your findings into data-driven recommendations for stakeholders, guiding the decision for a production-level rollout in a customer support application.

Das ist alles enthalten

2 Lektüren1 Aufgabe

Erwerben Sie ein Karrierezertifikat.

Fügen Sie dieses Zeugnis Ihrem LinkedIn-Profil, Lebenslauf oder CV hinzu. Teilen Sie sie in Social Media und in Ihrer Leistungsbeurteilung.

Dozent

Professionals from the Industry

472 Kurse83.884 Lernende

von

Coursera

Mehr von Machine Learning entdecken

Status: Kostenlos
DeepLearning.AI
Quality and Safety for LLM Applications
Projekt
Packt
LLM Engineer’s Handbook
Kurs
Status: Kostenloser Testzeitraum
Packt
Building and Fine-Tuning LLM Applications
Kurs
Status: Kostenloser Testzeitraum
Coursera
Optimize & Interface LLM Apps Effectively
Kurs

Warum entscheiden sich Menschen für Coursera für ihre Karriere?

Felipe M.

Lernender seit 2018

„Es ist eine großartige Erfahrung, in meinem eigenen Tempo zu lernen. Ich kann lernen, wenn ich Zeit und Nerven dazu habe.“

Jennifer J.

Lernender seit 2020

„Bei einem spannenden neuen Projekt konnte ich die neuen Kenntnisse und Kompetenzen aus den Kursen direkt bei der Arbeit anwenden.“

Larry W.

Lernender seit 2021

„Wenn mir Kurse zu Themen fehlen, die meine Universität nicht anbietet, ist Coursera mit die beste Alternative.“

Chaitanya A.

„Man lernt nicht nur, um bei der Arbeit besser zu werden. Es geht noch um viel mehr. Bei Coursera kann ich ohne Grenzen lernen.“

Häufig gestellte Fragen

The course assumes basic familiarity with statistics. It includes practical, applied lessons on confidence intervals and hypothesis testing, and offers step-by-step examples so that practitioners with modest statistical knowledge can follow along. Consider a short statistics refresher if you are new to hypothesis testing.

You will write evaluation scripts in Python, analyze logs and segmented datasets, run A/B test analyses, use SQL for data retrieval, and evaluate vector-search parameters (e.g., HNSW) commonly used with vector databases. No proprietary tools are required.

The course focuses on measurable, repeatable engineering practices: automated evaluation pipelines, statistical experiment design, log-driven debugging, and data-layer tuning. These skills help you prioritize fixes and validate improvements in real production settings.

To access the course materials, assignments and to earn a Certificate, you will need to purchase the Certificate experience when you enroll in a course. You can try a Free Trial instead, or apply for Financial Aid. The course may offer 'Full Course, No Certificate' instead. This option lets you see all course materials, submit required assessments, and get a final grade. This also means that you will not be able to purchase a Certificate experience.

Weitere Fragen

Besuchen Sie die das Hilfe-Center für Kursteilnehmer.

Finanzielle Unterstützung verfügbar,

¹ Einige Aufgaben in diesem Kurs werden mit AI bewertet. Für diese Aufgaben werden Ihre Daten in Übereinstimmung mit Datenschutzhinweis von Courseraverwendet.