The Economics of AI Inference GPT 4 class models cost $30/M tokens in 2024. Today they cost $3 — and DeepSeek V3 delivers comparable quality at $0.30. What's Driving the Drop 1. Custom inference chips (TPUs, Trainium, Groq LPUs) running 10 100x more efficiently than GPUs. 2. Open weights creating real competition. 3. Quantization — Q4 and Q5 models perform within 2% of full precision. Implications Margins on basic chat are gone. Value moves up the stack to agents, RAG, and vertical specialization.