What will I get if I subscribe to this Specialization?

When you enroll in the course, you get access to all of the courses in the Specialization, and you earn a certificate when you complete the work. Your electronic Certificate will be added to your Accomplishments page - from there, you can print your Certificate or add it to your LinkedIn profile.

Is financial aid available?

Yes. In select learning programs, you can apply for financial aid or a scholarship if you can’t afford the enrollment fee. If fin aid or scholarship is available for your learning program selection, you’ll find a link to apply on the description page.

Building Reliable LLM Systems

抓住节省的机会！购买 Coursera Plus 3 个月课程可享受40% 的折扣，并可完全访问数千门课程。

Building Reliable LLM Systems

本课程是 LLM Engineering That Works: Prompting, Tuning, and Retrieval 专项课程的一部分

位教师：Professionals from the Industry

包含在中

了解更多

5个模块

深入了解一个主题并学习基础知识。

中级等级

推荐体验

2 周完成

在 10 小时一周

灵活的计划

自行安排学习进度

5个模块

深入了解一个主题并学习基础知识。

中级等级

推荐体验

2 周完成

在 10 小时一周

灵活的计划

自行安排学习进度

您将学到什么

Build scripts with lexical/semantic metrics to evaluate LLMs, diagnose hallucinations, and balance vector-search recall against latency.
Apply hypothesis testing, confidence intervals, and significance metrics to evaluate model accuracy and validate results from A/B experiments.
Utilize parameterized SQL and data manipulation to segment user logs, calculate retention, and securely retrieve large-scale datasets.
Analyze LLM performance gaps to prioritize technical fixes and implement remediation measures for production-level reliability.

您将获得的技能

您将学习的工具

要了解的详细信息

可分享的证书

添加到您的领英档案

了解顶级公司的员工如何掌握热门技能

了解关于 Coursera for Business 的更多信息

Petrobras, TATA, Danone, Capgemini, P&G 和 L'Oreal 的徽标

积累特定领域的专业知识

本课程是 LLM Engineering That Works: Prompting, Tuning, and Retrieval 专项课程专项课程的一部分

在注册此课程时，您还会同时注册此专项课程。

向行业专家学习新概念
获得对主题或工具的基础理解
通过实践项目培养工作相关技能
获得可共享的职业证书

该课程共有5个模块

Building Reliable LLM Systems is a comprehensive course for AI practitioners looking to move beyond basic models and create production-grade applications. While getting an LLM to generate text is easy, ensuring a consistently accurate, relevant, and trustworthy output is a significant engineering challenge. This course provides a systematic framework for tackling the entire lifecycle of LLM reliability.

You will start by learning to quantitatively evaluate model performance using a suite of lexical and semantic metrics, such as BLEU, ROUGE-L, and cosine similarity. You’ll dive deep into debugging, using log analysis and data manipulation to uncover the root causes of critical failures, such as hallucinations, by correlating them with retrieval system performance. The course emphasizes statistical rigor, teaching you to design and analyze A/B tests, apply hypothesis testing, and calculate confidence intervals to prove the significance of your optimizations. Finally, you’ll optimize the foundational data layers, learning to tune SQL queries and vector search parameters to achieve the perfect balance between recall and latency.

This module lays the groundwork for quantitative Large Language Mode (LLM) evaluation. Learners will discover why relying on intuition to judge model performance is unsustainable and explore the foundational metrics used to create automated, objective evaluation systems. We will cover both lexical similarity metrics (like BLEU and ROUGE-L) that assess text structure and semantic metrics (like cosine similarity) that capture meaning. By the end of this module, learners will have the conceptual understanding and practical code to build their first automated evaluation script.

涵盖的内容

8个视频3篇阅读材料3个作业3个非评分实验室

8个视频总计44分钟

How to Compute Lexical Metrics: BLEU & ROUGE-L in Python? 6分钟
How to Compute Semantic Similarity with Embeddings? 6分钟
Why Guess When You Can Know? The Case of the "Better" Prompt 5分钟
The Language of Experimentation: Hypotheses, P-Values, and Power 5分钟
Running the Numbers: A/B Test Analysis in Python 7分钟
From Report to Action: The Optimization Loop 3分钟
Case Study: Benchmarking a Sentiment Analyzer 6分钟
Scripting Your First Evaluation Report 6分钟

3篇阅读材料总计17分钟

A Guide to LLM Evaluation: Lexical and Semantic Metrics 5分钟
Designing a Fair Race: A/B Testing for LLMs 7分钟
Building a Reproducible Evaluation Workflow 5分钟

3个作业总计50分钟

Build Your LLM Evaluation Toolkit 30分钟
Knowledge Check: Choosing Your Metrics 10分钟
Knowledge Check: Statistical Testing Concepts 10分钟

3个非评分实验室总计128分钟

Building Your First Automated Evaluation Script 60分钟
Statistical Significance Testing 60分钟
Planning Your Optimization Strategy 8分钟

When a production chatbot starts giving incorrect answers, how do you find the problem and fix it? This module equips AI practitioners, ML engineers, and data analysts with the essential skills for debugging production LLMs. Go beyond theory and learn the systematic, data-driven workflow that professionals use to solve the critical problem of AI hallucinations. You will be equipped to transition from merely observing AI failures to expertly diagnosing and resolving them.

涵盖的内容

5个视频3篇阅读材料3个作业2个非评分实验室

5个视频总计29分钟

Why Logs Matter: The Air Canada Case? 6分钟
Calculating Retention in Pandas 6分钟
Why RAG Fails: The Root of Hallucination? 6分钟
Correlating Errors with Retrieval in Pandas 6分钟
Visualizing the Proof in Matplotlib 5分钟

3篇阅读材料总计28分钟

Anatomy of a Log File 8分钟
The Engineering Brief: From Analysis to Action 10分钟
Authoring the Engineering Brief 10分钟

3个作业总计40分钟

LLM Diagnostics Report 30分钟
Knowledge Check: Retention Metrics 5分钟
Knowledge Check: Communicating Findings 5分钟

2个非评分实验室总计120分钟

Lab 1: Segmenting Users & Finding the Drop 60分钟
Lab 2: Proving the Root Cause 60分钟

When making high-stakes deployment decisions, a simple accuracy score is not enough. This module equips you with the statistical methods to rigorously validate LLM performance improvements. By the end of this module, you will be able to move beyond subjective "it seems better" evaluations to confidently state, "we can prove it's better," ensuring every deployment decision is backed by sound statistical evidence.

涵盖的内容

5个视频2篇阅读材料3个作业3个非评分实验室

5个视频总计30分钟

Why Single Scores Lie 8分钟
Calculating Wilson Intervals in Python 4分钟
Why Gut Feelings Fail in A/B Testing 6分钟
Running a Chi-Square Test in Python 6分钟
Visualizing Confidence with Matplotlib 5分钟

2篇阅读材料总计14分钟

Core Concepts: Confidence and Significance 8分钟
Storytelling with Statistical Visuals 6分钟

3个作业总计40分钟

LLM Evaluation Report 30分钟
Confidence Intervals Quiz 5分钟
Communicating Results Quiz 5分钟

3个非评分实验室总计110分钟

Lab 1: Quantifying Model Accuracy 20分钟
Lab 2: Validating a Model Improvement 30分钟
Lab 3: Create a Comparison Chart 60分钟

In the world of large-scale AI, slow queries and inefficient search can bring a system to its knees. This module provides the critical skills to prevent that, focusing on practical database and vector search optimization techniques. By the end of this module, you will be equipped to systematically analyze and optimize production retrieval systems, ensuring your AI applications are not only powerful but also fast and reliable.

涵盖的内容

4个视频3篇阅读材料4个作业3个非评分实验室

4个视频总计26分钟

From Inefficient to Optimized 7分钟
The Recall vs. Latency Trade-Off 5分钟
Tuning an HNSW Index 8分钟
Beyond One-Off Tests: The Need for Continuous Benchmarking 5分钟

3篇阅读材料总计25分钟

Secure and Efficient Query Patterns 10分钟
Understanding Vector Search Parameters 10分钟
Core Metrics of a Benchmarking Framework 5分钟

4个作业总计85分钟

Submit Your Performance Optimization Report 45分钟
SQL Security and Patterns 15分钟
Parameter Tuning Scenarios Quiz 15分钟
Interpreting Benchmark Results 10分钟

3个非评分实验室总计140分钟

Identifying Slowest Queries using Parameterized SQL 20分钟
Tune HNSW Parameters for Recall and Latency 60分钟
Create an Automated Benchmarking Suite 60分钟

In this module, you will conduct an end-to-end performance audit comparing two LLM variants using an A/B test dataset. You will implement a pipeline to calculate key performance metrics, including lexical and semantic similarity, and use statistical A/B testing to validate model improvements. The project culminates in a comprehensive report where you will correlate hallucination rates with retrieval logs and synthesize your findings into data-driven recommendations for stakeholders, guiding the decision for a production-level rollout in a customer support application.

涵盖的内容

2篇阅读材料1个作业

获得职业证书

将此证书添加到您的 LinkedIn 个人资料、简历或履历中。在社交媒体和绩效考核中分享。

位教师

Professionals from the Industry

217 门课程 34,516 名学生

提供方

Coursera

从 Machine Learning 浏览更多内容

状态：免费
DeepLearning.AI
Quality and Safety for LLM Applications
项目
Packt
LLM Engineer’s Handbook
课程
状态：免费试用
Packt
Building and Fine-Tuning LLM Applications
课程
状态：免费试用
Coursera
Optimize & Interface LLM Apps Effectively
课程

人们为什么选择 Coursera 来帮助自己实现职业发展

Felipe M.

自 2018开始学习的学生

''能够按照自己的速度和节奏学习课程是一次很棒的经历。只要符合自己的时间表和心情，我就可以学习。'

Jennifer J.

自 2020开始学习的学生

''我直接将从课程中学到的概念和技能应用到一个令人兴奋的新工作项目中。'

Larry W.

自 2021开始学习的学生

''如果我的大学不提供我需要的主题课程，Coursera 便是最好的去处之一。'

Chaitanya A.

''学习不仅仅是在工作中做的更好：它远不止于此。Coursera 让我无限制地学习。'

通过 Coursera Plus 开启新生涯

无限制访问 10,000+ 世界一流的课程、实践项目和就业就绪证书课程 - 所有这些都包含在您的订阅中

了解更多

通过在线学位推动您的职业生涯

获取世界一流大学的学位 - 100% 在线

探索学位

加入超过 3400 家选择 Coursera for Business 的全球公司

提升员工的技能，使其在数字经济中脱颖而出

了解更多

常见问题

The course assumes basic familiarity with statistics. It includes practical, applied lessons on confidence intervals and hypothesis testing, and offers step-by-step examples so that practitioners with modest statistical knowledge can follow along. Consider a short statistics refresher if you are new to hypothesis testing.

You will write evaluation scripts in Python, analyze logs and segmented datasets, run A/B test analyses, use SQL for data retrieval, and evaluate vector-search parameters (e.g., HNSW) commonly used with vector databases. No proprietary tools are required.

The course focuses on measurable, repeatable engineering practices: automated evaluation pipelines, statistical experiment design, log-driven debugging, and data-layer tuning. These skills help you prioritize fixes and validate improvements in real production settings.

To access the course materials, assignments and to earn a Certificate, you will need to purchase the Certificate experience when you enroll in a course. You can try a Free Trial instead, or apply for Financial Aid. The course may offer 'Full Course, No Certificate' instead. This option lets you see all course materials, submit required assessments, and get a final grade. This also means that you will not be able to purchase a Certificate experience.