High-Throughput Edge AI & Model Quantization
4-bit AWQ and GGUF quantization, TensorRT-LLM compilation, zero-latency inference scaling, and embedded edge deployments.
Course Overview & Objectives
Deploy massive models on resource-constrained hardware with zero compromise on safety. Master 4-bit AWQ, GPTQ, and GGUF quantization, compile optimized TensorRT-LLM inference engines, and run models locally on embedded edge devices.
What You Will Master
- Quantize 70B parameter models down to 4-bit INT4 precision with minimal perplexity degradation
- Build and benchmark high-throughput inference servers with vLLM continuous batching
- Compile custom TensorRT-LLM engines for NVIDIA Jetson and desktop GPUs
- Audit quantized models for emerging safety alignment degradation flaws
Prerequisites
- Python, Linux CLI, and basic understanding of GPU hardware architectures
Platforms & Tools Covered
Detailed Curriculum Modules
1 modules structured from foundational theory through complex adversarial execution.
Precision Formats & Quantization Mechanics
FP32, FP16, BF16, INT8, and INT4 arithmetic, outlier weights, and perplexity evaluations.
Hands-on Virtual Sandbox Labs
Zero local hardware dependencies. Provisioned in cloud containers via browser terminal.
Deploying High-Throughput vLLM Cluster with PagedAttention
Benchmark tokens-per-second throughput under 100 concurrent user streams.
Faculty & Lead Instructor
Direct weekly instruction, live office hours, and code-review feedback.
Harpreet Kaur
Thread Security EducationSenior AI Engineer
Optimizing high-concurrency LLM inference fleets and embedded edge deployments for real-time defense applications.
Frequently Asked Questions
Everything you need to know about scheduling, cohort admissions, and lab access.
Can I run these quantized models on a standard laptop?
Yes! The course teaches both cloud GPU clusters and local GGUF/Ollama deployments.
Ready to Master High-Throughput Edge AI & Model Quantization?
Join the upcoming cohort. Seats are limited to maintain a high faculty-to-student ratio and rigorous sandbox feedback.