Skip to content

Add comprehensive documentation and 24-session learning curriculum - #82

Open
Kandil7 wants to merge 10 commits into
PacktPublishing:mainfrom
Kandil7:main
Open

Kandil7 wants to merge 10 commits into
PacktPublishing:mainfrom
Kandil7:main

Conversation

@Kandil7

@Kandil7 Kandil7 commented Oct 3, 2026

Copy link
Copy Markdown

Summary

Adds a complete documentation set and a 24-session learning curriculum for the LLM Engineer's Handbook, written against the actual source code in this repository.

What's included

  • Curriculum and guides: docs/CURRICULUM.md, docs/GETTING_STARTED.md, and docs/README.md (documentation index).
  • Code intelligence: docs/CODEBASE-INTELLIGENCE.md with architecture, graphs, and tooling references.
  • 24 session walkthroughs under docs/sessions/, covering:
    • Foundations: project overview, domain layer (ODM), infrastructure layer (MongoDB/Qdrant).
    • Data engineering: web crawling, text preprocessing, feature engineering.
    • Datasets: instruction dataset generation, preference dataset for DPO.
    • RAG: advanced RAG architecture, embedding models and cross-encoders.
    • Training: SFT/QLoRA, DPO, AWS SageMaker deployment.
    • Inference: FastAPI service, RAG inference flow.
    • Monitoring: Comet ML experiment tracking, Opik tracing, model evaluation.
    • Production: Docker, CI/CD with GitHub Actions, ZenML orchestration.
    • Advanced: data warehouse backup/restore, performance tuning, security hardening.
  • .gitignore: ignore the local graphify-out/ artifacts directory.

Structure

Each session follows the same format: learning objectives, architecture overview, per-file deep dives with code, hands-on steps, an exercise, a knowledge check, and next steps.

Notes

  • This is a documentation-only change; no runtime code is modified.
  • Content is grounded in the real modules, class names, and configuration values in the repository.

Kandil7 added 10 commits October 3, 2026 21:25
- Created session 2.1 on Web Crawling with Selenium, covering architecture, key files, error handling, and best practices.
- Introduced session 4.1 on Advanced RAG Architecture, detailing multi-stage retrieval, query expansion, reranking, and self-query techniques.
- Introduced Session 2.2: Text Preprocessing Pipeline, detailing the architecture, key files, and hands-on exercises for cleaning and chunking documents.
- Added Session 2.3: Feature Engineering Pipeline, covering embedding models, batch processing, and loading embedded documents into Qdrant.
- Explained the Singleton pattern for thread-safe model instantiation and the embedding process for various document types.
- Included exercises and knowledge checks to reinforce learning objectives.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant