https://i128.fastpic.org/big/2026/0924/a4/eb41874812babb8eb0d43cbf5b1374a4.jpg
(English)RAG with Local LLMs: Ollama, LangChain &; ChromaDB
Published 9/2026
MP4 | Video: h264, 1920x1080 | Audio: AAC, 44.1 KHz, 2 Ch
Language: English + subtitle | Duration: 1h 40m | Size: 280.9 MB

Build document Q&A in Python, evaluate retrieval and answers, and improve a local RAG system without cloud LLM APIs.

What you'll learn
Run local generation and embedding models with Ollama and call them from Python.
Explain sentence embeddings and calculate cosine similarity with NumPy.
Extract, clean and split text while preserving context and source metadata.
Build and update a persistent ChromaDB document index.
Compose a retriever, prompt, local LLM and parser with LangChain and LCEL.
Build an evaluation set and distinguish retrieval failures from generation failures.
Compare BM25, hybrid retrieval, RRF, reranking and prompt changes using recorded results.
Create a local Gradio document QA app with sources, updates and follow-up questions.

Requirements
Basic Python: variables, functions, imports and running a script.
Comfort with a terminal and installing development tools.
Python 3.12, Ollama and a computer suitable for local models; 16 GB RAM or more is recommended.
Allow at least 15 GB of free disk space, with additional space for optional models and your own documents.
An internet connection for initial package and model downloads. No paid cloud LLM API key is required.

Description
This course includes the use of AI.

English narration and English captions are included. This course uses AI-generated narration and AI-assisted translation. The instructor is responsible for the course content. Turn on captions to follow technical terms and code more easily.

Build a document question-answering system that runs on your own computer, then learn how to test whether it actually answers well.

A working RAG demo is only the beginning. The more useful question is: when it gives the wrong answer, can you find out why? Was the evidence missing from retrieval? Did chunking separate a condition from its exception? Did the prompt make the model refuse a question it could answer? Or did a more expensive retrieval technique improve a ranking metric without improving the final answer?

In this practical course, you will build the complete workflow with Python, Ollama, Qwen2.5, bge-m3, ChromaDB and LangChain. After downloading the required models and packages, document retrieval and answer generation use local models. No paid cloud LLM API key is required for the course application. Initial installation and optional model downloads require an internet connection.

The course is organized into eight sections and six implementation phases. Start with setup, then explore embeddings and cosine similarity with NumPy. Extract and clean text, design chunks and metadata, store vectors in ChromaDB, and connect retrieval to generation using LangChain and LCEL. Finally, build a local Gradio application with source display, incremental indexing and follow-up questions.

Evaluation is central. You will work with a 30-question English evaluation set covering direct questions, identifiers, negation, questions spanning documents and paraphrases. Measure Recall@k, MRR and nDCG, inspect answer checks and refusals, and compare prompt variants, BM25, hybrid retrieval, Reciprocal Rank Fusion and reranking. An oracle experiment helps separate retrieval problems from generation problems. Keyword checks are a practical debugging aid, not proof of semantic correctness; the course also emphasizes inspecting the actual answers and their evidence.

Downloads include English notebooks, executable code, fictional corporate policies, teaching PDFs, evaluation questions and copyable lecture snippets. The fictional policies retain their original case-study currency and conditions. They are learning data, not employment, legal or compliance advice.

The examples use a local development environment with Python 3.12. A machine with at least 16 GB of RAM and sufficient free disk space is recommended; smaller machines may need smaller models and will have different performance. The setup examples focus on macOS, with Windows activation and startup commands supplied in the downloads. Runtime behavior and evaluation results can change with model and package versions, hardware and documents.

You should already understand basic Python functions, imports and terminal commands. You do not need previous RAG experience. This course is for developers who want to understand the system they are building, diagnose its failures, and make measured decisions before adding more complexity.

By the end, you will have a local document QA application and a repeatable process for evaluating changes on your own authorized documents. Local processing reduces dependence on cloud services, but it does not replace access controls, source verification or careful handling of confidential information.

Who this course is for
Python developers building document search or question-answering prototypes.
Engineers who want to evaluate and debug RAG rather than stop at a basic demo.
Developers exploring local models for workflows where document handling needs careful control.
Learners who know basic Python and want a practical introduction to embeddings, vector databases and grounded generation.

Homepage

Код:
https://www.udemy.com/course/englishrag-with-local-llms-ollama-langchain-chromadb/

https://rapidgator.net/file/9ac1bca3072871dfc337bef3ee633af6/(English)RAG_with_Local_LLMs_Ollama,_LangChain_&_ChromaDB.rar.html

https://www.uploadcloud.pro/75xv6le6a3n … B.rar.html