Writing

All Articles

29 deep dives on architectures, research, and the ideas shaping modern AI.

Fine-Tuning & LLMs Sep 24, 2026

Fine-Tuning Open-Weights Models with QLoRA & Unsloth

A practical hands-on guide to parameter-efficient fine-tuning (PEFT): 4-bit quantization, Low-Rank Adaptation (LoRA), Memory footprint management, and 2x faster training with Unsloth.

Read article
LLM & Agents Sep 24, 2026

Multi-Agent Orchestration: Designing Reliable Autonomous Workflow Systems

A comprehensive technical architectural guide to building production multi-agent systems: task decomposition, agent delegation, state isolation, handoff contracts, and failure recovery.

Read article
RAG & Retrieval Sep 24, 2026

Production RAG Engineering: Hybrid Search, Reranking & Evaluation

An end-to-end technical deep dive into moving beyond naive vector RAG: dense+sparse hybrid search, cross-encoder reranking, chunking strategies, and automated evaluation metrics.

Read article
Reasoning & Research Sep 24, 2026

Test-Time Compute & Inference Scaling for Reasoning Models

A deep technical analysis of the shift from pre-training compute scaling to inference-time scaling: Monte Carlo Tree Search (MCTS), Process Reward Models (PRMs), and self-correction loops.

Read article
Agents Sep 06, 2026

Agent Harnesses: The Engineering System Around Reliable AI Agents

A technical guide to building agent harnesses: the runtime, state, tools, policies, evaluation, and observability that turn an LLM loop into a dependable system.

Read article
MLOps Sep 06, 2026

Deploying Models in Production: A Technical Guide from Artifact to SLO

A practical deep dive into production model deployment: packaging, serving, scaling, rollout strategies, observability, and the failure modes that matter after training.

Read article
LLM Architecture Sep 06, 2026

Kimi Delta Attention: A Technical Deep Dive into Gated Linear Memory

How Kimi Delta Attention replaces quadratic attention with a gated delta-rule memory, and what that means for long-context language models.

Read article
LLM & Agents Aug 30, 2026

Agentic AI and the Model Context Protocol (MCP)

A senior engineer's tour of the path from plain RAG to autonomous, multi-agent systems, and how the Model Context Protocol standardizes the way LLM apps connect to tools and data.

Read article
Research Aug 16, 2026

Reasoning Models and Test-Time Compute

How the field shifted from scaling training to spending compute at inference — chain-of-thought as computation, RLVR, and the sampling, verification, and search strategies that turn extra thinking into better answers.

Read article
Deep Learning Jul 20, 2026

Mamba and State Space Models: The Transformer Challenger

A senior-level tour of state space models — from S4 to Mamba's selective SSMs and Mamba-2's state-space duality — as a linear-time alternative to quadratic self-attention.

Read article
Deep Learning Jun 28, 2026

Mixture-of-Experts: How Frontier LLMs Scale Efficiently

A senior-level walkthrough of Mixture-of-Experts in LLMs — how sparse routing decouples parameter count from per-token compute, plus the losses, failure modes, and real systems that make it work.

Read article
Research May 25, 2026

Foundation Models for Causal Inference (CausalFM)

CausalFM is a transformer-based foundation model trained on simulated causal worlds so it can reason about cause and effect, estimate treatment effects, and answer counterfactual what-if questions rather than just detect correlations.

Read article
Vision-Language Jan 12, 2026

VL JEPA: A Deep Dive into Vision-Language Predictive Models

A detailed engineering walkthrough of VL JEPA, the Vision-Language Joint Embedding Predictive Architecture, situated within Meta AI's JEPA lineage (I-JEPA, V-JEPA, V-JEPA 2) and the broader shift toward predictive, non-generative representation learning.

Read article
Computer Vision Dec 20, 2025

Understanding the CLIP Model: Bridging Vision and Language

A deep, practical walkthrough of OpenAI's CLIP: the dual-encoder architecture, the contrastive (InfoNCE) training objective, prompt engineering, worked zero-shot classification, and how CLIP-style encoders became the backbone of modern multimodal systems through 2025-2026.

Read article
Research Dec 11, 2025

Titans Architecture and the MIRAS Framework

A deep dive into Google Research's Titans architecture and the MIRAS framework: how test-time memorization, surprise-driven updates, and attentional bias push models past the quadratic-attention wall.

Read article
Deep Learning Dec 11, 2025

Transformers++: The Next Evolution in AI Architecture

A deep, engineering-level tour of Transformers++ — the wave of architectural changes (linear and sparse attention, state-space hybrids, neural memory, mixture-of-experts, and efficient KV-cache serving) that extend the classic Transformer toward long context, multimodality, and continuous learning.

Read article
Deep Learning Dec 11, 2025

Transformers: The Architecture That Changed AI Forever

A deep, modern dive into Transformer architecture — attention math, encoder/decoder design, the 2025-2026 stack (RoPE, RMSNorm, GQA, FlashAttention, MoE, long context), efficiency techniques, evaluation, and where the field is heading.

Read article
Personal Nov 24, 2025

800 Days of Grind: The Full Blueprint

An 800-day blueprint for transformation, backed by the neuroscience of habit formation, identity-based behavior change, deliberate practice, and a green/yellow/red tracking system that survives real life.

Read article
Neuroscience Nov 24, 2025

How the Human Mind Works

A deep dive into how the human mind works — the brain's architecture, conscious and subconscious layers, perception, memory, emotion, creativity, and what modern neuroscience reveals about cognition and consciousness.

Read article
LLM Optimization Oct 02, 2025

KV Cache Compression & Context Window Extension: YaRN, RoPE & PagedAttention

Managing gigabyte KV caches for 1M+ token contexts: vLLM PagedAttention, Rotary Position Embeddings (RoPE), YaRN frequency scaling, and Grouped-Query Attention.

Read article
Distributed Training Sep 14, 2025

Distributed Training at Scale: Megatron-LM 3D Parallelism & DeepSpeed ZeRO-3

Scaling LLM training across 10,000+ GPUs: Tensor Parallelism (TP), Pipeline Parallelism (PP), Sequence Parallelism (SP), and ZeRO-3 Memory Offloading.

Read article
LLM Architecture Aug 23, 2025

Mixture of Depths (MoD): Dynamic Layer Routing & Compute Allocation

Routing compute dynamically per token across transformer layers: top-k routing, capacity bounds, and integrating MoD with Mixture-of-Experts (MoE).

Read article
AI Safety & Interpretability Jul 12, 2025

Mechanistic Interpretability & Sparse Autoencoders: Deconstructing LLM Polysemanticity

Extracting monosemantic features from LLM activations using dictionary learning, Sparse Autoencoders (SAEs), and circuit probing.

Read article
Generative Modeling Jun 22, 2025

Diffusion Models & Flow Matching: Continuous-Time Generative Architectures

The mathematical formulation of Score-based Generative Models, SDE/ODE formulations, Rectified Flow Matching, and Classifier-Free Guidance (CFG).

Read article
Reinforcement Learning May 05, 2025

RLVR & Reasoning Alignment: Reinforcement Learning with Verifiable Rewards

How RLVR replaces subjective reward models with objective code compilers and math solvers: Group Relative Policy Optimization (GRPO) and self-play.

Read article
Inference Optimization Apr 19, 2025

Speculative Decoding & Medusa Heads: Accelerating LLM Token Generation

Scaling LLM inference throughput: draft-target model speculative sampling, tree-based verification, and multi-head parallel token prediction with Medusa.

Read article
Multimodal AI Mar 22, 2025

Vision-Language Multimodal Architecture: SigLIP, LLaVA & Patch Tokenization

Engineering vision-language models: ViT patch embeddings, SigLIP sigmoid loss vs InfoNCE, linear projection layers, and cross-attention spatial grounding.

Read article
LLM Alignment Feb 22, 2025

Direct Preference Optimization (DPO) & Alignment: Beyond RLHF with PPO

A rigorous mathematical breakdown of LLM alignment techniques: Bradley-Terry preference modeling, PPO loss dynamics, DPO implicit reward derivation, KTO, and ORPO.

Read article
GPU Systems & Kernels Jan 25, 2025

FlashAttention-3 & GPU Kernel Optimization: Accelerating Transformer Attention

An in-depth technical analysis of FlashAttention-1, 2, and 3: IO-awareness, FP8 Tensor Cores, async memory pipelines, and custom Triton kernel design.

Read article