Production AI InfrastructureExpert Level70 Hours Live

High-Throughput Edge AI & Model Quantization

4-bit AWQ and GGUF quantization, TensorRT-LLM compilation, zero-latency inference scaling, and embedded edge deployments.

AWQ & GGUF 4-Bit Precision Quantization
TensorRT-LLM Engine Optimization
vLLM Continuous Batching Architecture
70 Hours Practical Workload
1 Core Modules
1 Sandboxed Labs
Cryptographic TS-ID Verifiable

Course Overview & Objectives

Deploy massive models on resource-constrained hardware with zero compromise on safety. Master 4-bit AWQ, GPTQ, and GGUF quantization, compile optimized TensorRT-LLM inference engines, and run models locally on embedded edge devices.

What You Will Master

  • Quantize 70B parameter models down to 4-bit INT4 precision with minimal perplexity degradation
  • Build and benchmark high-throughput inference servers with vLLM continuous batching
  • Compile custom TensorRT-LLM engines for NVIDIA Jetson and desktop GPUs
  • Audit quantized models for emerging safety alignment degradation flaws

Prerequisites

  • Python, Linux CLI, and basic understanding of GPU hardware architectures

Platforms & Tools Covered

vLLMTensorRT-LLMAutoAWQllama.cppOllama

Detailed Curriculum Modules

1 modules structured from foundational theory through complex adversarial execution.

70 Total Workload Hours
MODULE 01

Precision Formats & Quantization Mechanics

1 Lessons

FP32, FP16, BF16, INT8, and INT4 arithmetic, outlier weights, and perplexity evaluations.

Quantizing Llama 3 with AutoAWQ and Evaluating Accuracy
50m

Hands-on Virtual Sandbox Labs

Zero local hardware dependencies. Provisioned in cloud containers via browser terminal.

LAB 01~60 mins

Deploying High-Throughput vLLM Cluster with PagedAttention

Benchmark tokens-per-second throughput under 100 concurrent user streams.

Skills Tested:vLLM, Inference Optimization

Faculty & Lead Instructor

Direct weekly instruction, live office hours, and code-review feedback.

HK

Harpreet Kaur

Thread Security Education

Senior AI Engineer

Optimizing high-concurrency LLM inference fleets and embedded edge deployments for real-time defense applications.

Frequently Asked Questions

Everything you need to know about scheduling, cohort admissions, and lab access.

Can I run these quantized models on a standard laptop?

Yes! The course teaches both cloud GPU clusters and local GGUF/Ollama deployments.

Ready to Master High-Throughput Edge AI & Model Quantization?

Join the upcoming cohort. Seats are limited to maintain a high faculty-to-student ratio and rigorous sandbox feedback.