Complete AI & Machine Learning Mind Map (2026)
Tip: the map is wide — swipe / scroll horizontally to explore, or tap a highlighted node.
Nodes with a dashed teal border are clickable — jump to a diagram of how that algorithm works.
Jump to algorithm diagrams ↓
graph LR
ROOT(("Artificial Intelligence & Machine Learning")):::root
%% ============ LEFT SIDE (branches flow into the center) ============
%% Supervised
CLS["Classification • Logistic Regression: $P(y=1|x) = \frac{1}{1 + e^{-(\beta_0 + \beta_1 x)}}$ • SVM: Maximize margin $2 / ||w||$ • Decision Trees: Entropy = $-\sum p_i \log p_i$ • Random Forest: Bagging + Trees • KNN: Euclidean Distance $\sqrt{\sum (x_i - y_i)^2}$ • Naive Bayes: $P(y|x) = \frac{P(x|y)P(y)}{P(x)}$ • XGBoost: Gradient Boosting + Trees • LightGBM: Leaf-wise Trees • CatBoost: Categorical Handling • Neural Networks: Softmax Output"]:::sub
REG_S["Regression • Linear: $y = \beta_0 + \beta_1 x + \epsilon$ • Polynomial: $y = \beta_0 + \beta_1 x + \beta_2 x^2 + ...$ • Ridge: L2 Reg. $\lambda \sum \beta_i^2$ • Lasso: L1 Reg. $\lambda \sum |\beta_i|$ • ElasticNet: L1 + L2 Combo • SVR: Epsilon-Tube Margin • Gradient Boosting: Sequential Trees"]:::sub
CLS --- SUP["Supervised Learning"]:::main
REG_S --- SUP
SUP --- ROOT
%% Unsupervised
CLU["Clustering • K-Means: Minimize $\sum ||x_i - \mu_j||^2$ • Hierarchical: Dendrogram Distance • DBSCAN: Density-Based $\epsilon$-Neighborhood • OPTICS: Variable Density • Gaussian Mixture: EM Algo $P(x) = \sum \pi_k N(x|\mu_k, \Sigma_k)$ • Mean Shift: Kernel Bandwidth"]:::sub
DIM["Dimensionality Reduction • PCA: Eigen Decomposition $\Sigma = U \Lambda U^T$ • Kernel PCA: Non-Linear Mapping • t-SNE: KL Divergence Min. • UMAP: Manifold Approximation • Autoencoders: Reconstruction Loss • LDA: Class Separation • ICA: Independent Components"]:::sub
CLU --- UNS["Unsupervised Learning"]:::main
DIM --- UNS
UNS --- ROOT
%% Semi-Supervised
SEMI["Semi-Supervised Learning • Self-Training: Iterative Labeling • Label Propagation: Graph Diffusion • Co-Training: Multi-View Features"]:::sub
SEMI --- ROOT
%% Reinforcement Learning
RL["Reinforcement Learning & Alignment • Q-Learning: $Q(s,a) = r + \gamma \max Q(s',a')$ • SARSA: On-Policy Update • DQN / DDQN: Neural Q-Approx. • PPO: Clipped Surrogate • RLVR / GRPO: Verifiable Rewards • A2C / A3C: Advantage Actor-Critic • SAC: Entropy Regularized"]:::sub
RL --- ROOT
%% Deep Learning
ANN["ANN / MLP • Feedforward: $h = \sigma(Wx + b)$ • Backpropagation: Chain Rule $\frac{\partial L}{\partial w}$"]:::sub
CNN["CNN • Convolution: $ (f * g)(i,j) = \sum f(m,n) g(i-m,j-n) $ • Pooling: Max/Avg • ResNet: Skip Connections • YOLO: Bounding Boxes • U-Net: Encoder-Decoder • EfficientNet: Compound Scaling • MobileNet: Depthwise Conv."]:::sub
RNN["RNN • Vanilla: $h_t = \tanh(W h_{t-1} + U x_t)$ • LSTM: Gates (Forget/Input/Output) • GRU: Update/Reset Gates • Bi-directional: Forward + Backward • Seq2Seq: Encoder-Decoder"]:::sub
TRANS["Transformers & Next-Gen Architectures • Attention: $ \text{Attn} = \text{softmax}(QK^T / \sqrt{d}) V $ • FlashAttention-3: FP8 & Async SRAM • Mamba-2 & SSMs: State Space Duals • Mixture of Depths (MoD): Dynamic Compute • Mixture of Experts (MoE): Top-k Routing • Speculative Decoding & Medusa • ViT & SigLIP: Vision Transformers"]:::sub
LLM["Reasoning Models & LLMs • DeepSeek R1 & o1: Test-Time Compute • Gemini 1.5 Pro: 2M Context • Claude 3.5 Sonnet: Computer Use • Llama 3 405B: Open Weights • Qwen 2.5: Reasoning & Code"]:::sub
GEN["Generative Models • Flow Matching: Continuous ODEs • Diffusion: Denoising $p(x_{t-1}|x_t)$ • Flux.1 & SD3: Rectified Flow • GANs: Min-Max $V(G,D)$ • DALL·E 3: CLIP Guided • Sora: Video Generation"]:::sub
GNN["Graph Neural Networks • Message Passing • GCN: Spectral Convolution • GAT: Attention on Graphs • GraphSAGE: Neighbor Sampling • GIN: Expressive Power"]:::sub
ANN --- DL["Deep Learning"]:::main
CNN --- DL
RNN --- DL
TRANS --- DL
LLM --- DL
GEN --- DL
GNN --- DL
DL --- ROOT
%% LLM Engineering & GenAI Ops
RAG["Production RAG & Vector DBs • Hybrid Search: BM25 + Dense • Reciprocal Rank Fusion (RRF) • Cross-Encoder Reranking • Vector DB: FAISS / Pinecone / Qdrant • Chunking & Context Compression"]:::sub
AGENT["Agentic AI & MCP Systems • Model Context Protocol (MCP) • ReAct & Subagent Orchestration • Function Calling & Tool Isolation • Stateful Workflows & Memory Banks • A2A Protocol & Multi-Agent Swarms"]:::sub
PROMPT["Prompt & Reasoning Engineering • Process Reward Models (PRM) • Chain-of-Thought (CoT) Verification • Tree-of-Thoughts / MCTS • Structured Output & JSON Schema • Self-Consistency Sampling"]:::sub
FTUNE["Fine-Tuning, Alignment & Systems • RLVR / GRPO: Verifiable Rewards • DPO / KTO / ORPO: Direct Preference • QLoRA / Unsloth 4-bit NF4 • Triton CUDA Kernels • 3D Parallelism & ZeRO-3"]:::sub
LLMEVAL["Observability & Evaluation • LLM-as-Judge & Quality Flywheel • RAGAS & Groundedness Metrics • Sparse Autoencoders (SAEs) • Safety, Red Teaming & Guardrails"]:::sub
RAG --- LLMENG["LLM Engineering"]:::main
AGENT --- LLMENG
PROMPT --- LLMENG
FTUNE --- LLMENG
LLMEVAL --- LLMENG
LLMENG --- ROOT
%% ============ RIGHT SIDE (branches flow out from the center) ============
%% Domains
CV["Computer Vision • Object Detection: IoU • Segmentation: Dice Coeff. • Pose Estimation: Keypoints • OCR: CTC Loss"]:::sub
NLP["Natural Language Processing • NER: Entity F1 • Sentiment: Polarity Scores • Translation: BLEU • Summarization: ROUGE"]:::sub
AUDIO["Audio & Speech • ASR: WER Metric • TTS: MOS Score • Whisper: CTC • Wav2Vec: Self-Supervised • HuBERT: Clustering • VALL-E: Codec • AudioLM: Tokenization"]:::sub
MULTI["Multimodal • CLIP: Contrastive Loss • LLaVA: Vision-Language • Flamingo: Few-Shot • ImageBind: Bind Modalities • Kosmos: Unified"]:::sub
TS["Time Series & Forecasting • ARIMA / SARIMA • Exponential Smoothing • Prophet • DeepAR / N-BEATS • Temporal Fusion Transformer"]:::sub
RECSYS["Recommender Systems • Collaborative Filtering • Matrix Factorization: SVD / ALS • Content-Based Filtering • Factorization Machines • Two-Tower / Neural CF"]:::sub
ROOT --- DOM["Domains"]:::main
DOM --- CV
DOM --- NLP
DOM --- AUDIO
DOM --- MULTI
DOM --- TS
DOM --- RECSYS
%% Core Components
OPT["Optimizers • SGD: $\theta = \theta - \eta \nabla L$ • Adam: Bias-Corrected Moments • AdamW: Decoupled Decay • RMSprop: Adaptive LR • Lion: Sign Momentum • LAMB: Layer-Wise"]:::sub
LOSS["Loss Functions • Cross-Entropy: $-\sum y \log \hat{y}$ • MSE: $\frac{1}{n} \sum (y - \hat{y})^2$ • Focal: Modulated CE • Contrastive: NT-Xent • CTC: Alignment"]:::sub
REGU["Regularization • Dropout: Random Mask • BatchNorm: $\hat{x} = \gamma (x - \mu)/\sigma + \beta$ • LayerNorm: Per-Feature • Weight Decay: L2 on Weights • Augmentation: Random Transforms"]:::sub
ROOT --- CORE["Core Components"]:::main
CORE --- OPT
CORE --- LOSS
CORE --- REGU
%% Advanced Techniques
ENS["Ensemble Methods • Bagging: Bootstrap Agg. • Boosting: Weighted Errors • XGBoost: 2nd-Order Grad. • LightGBM: GOSS • CatBoost: Ordered Boost."]:::sub
TRANSF["Transfer Learning • Fine-Tuning: Last Layers • Prompt Tuning: Soft Prompts • Few-Shot: In-Context • Zero-Shot: Pre-Trained"]:::sub
PROBM["Probabilistic Models • Gaussian Processes • Hidden Markov Models • Bayesian Networks • MCMC / Variational Inference • Kalman Filters"]:::sub
ROOT --- ADV["Advanced Techniques"]:::main
ADV --- ENS
ADV --- TRANSF
ADV --- PROBM
%% Foundations
LA["Linear Algebra • Vectors: Dot Product $u \cdot v$ • Matrices: Det(A) • SVD: $A = U \Sigma V^T$ • Eigen: $A v = \lambda v$"]:::sub
CALC["Calculus • Derivatives: $f'(x)$ • Gradients: $\nabla f$ • Chain Rule: $(f \circ g)' = f' g'$"]:::sub
PROB["Probability & Stats • Distributions: PDF/PMF • Bayes: $P(A|B) = P(B|A)P(A)/P(B)$ • MLE: Argmax Log-Lik."]:::sub
EVAL["Evaluation • Accuracy / F1: $2PR/(P+R)$ • ROC-AUC: TPR vs FPR • BLEU / ROUGE: N-Gram • FID / IS: Generative • Cross-Val: K-Fold"]:::sub
ROOT --- FOUND["Foundations"]:::main
FOUND --- MATH["Mathematics"]:::main
MATH --- LA
MATH --- CALC
MATH --- PROB
FOUND --- EVAL
%% Frameworks & Tools
SK["Scikit-learn • Pipelines • Grid Search"]:::sub
TORCH["PyTorch • Dynamic Graphs • Lightning: Trainer • TorchVision: Datasets"]:::sub
TF["TensorFlow • Static Graphs • Keras: Sequential • TF Hub: Pre-Trained"]:::sub
HF["Hugging Face • Transformers Lib. • Datasets Hub"]:::sub
TOOLS["Others • JAX: Autodiff • Gradio: UI • Streamlit: Apps • OpenCV: CV • W&B: Logging"]:::sub
ROOT --- FRAME["Frameworks & Tools"]:::main
FRAME --- SK
FRAME --- TORCH
FRAME --- TF
FRAME --- HF
FRAME --- TOOLS
%% MLOps & Deployment
SERVE["Serving & Inference • REST / gRPC APIs • TorchServe / Triton • vLLM / TGI (LLM serving) • Batch vs Real-Time"]:::sub
OPTZ["Model Optimization • Quantization: INT8 / 4-bit • Distillation: Teacher → Student • Pruning • ONNX / TensorRT"]:::sub
MONITOR["Monitoring & Lifecycle • Data / Concept Drift • Tracking: W&B / MLflow • CI/CD + Model Registry • Feature Store"]:::sub
ROOT --- MLOPS["MLOps & Deployment"]:::main
MLOPS --- SERVE
MLOPS --- OPTZ
MLOPS --- MONITOR
%% Responsible AI
XAI["Explainability (XAI) • SHAP: Shapley Values • LIME: Local Surrogates • Grad-CAM: Saliency Maps • Integrated Gradients"]:::sub
FAIR["Fairness & Ethics • Bias Metrics • Demographic Parity • Equalized Odds • Data / Label Bias"]:::sub
PRIV["Privacy & Robustness • Differential Privacy • Federated Learning • Adversarial Robustness • Model / Data Security"]:::sub
ROOT --- RAI["Responsible AI"]:::main
RAI --- XAI
RAI --- FAIR
RAI --- PRIV
classDef root fill:#0f766e,stroke:#134e4a,stroke-width:4px,color:white,font-weight:bold,font-size:20px
classDef main fill:#15803d,stroke:#166534,stroke-width:3px,color:white,font-weight:bold
classDef sub fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:black,font-weight:500
%% Clickable nodes -> algorithm diagrams (adds a dashed teal accent)
classDef clickable stroke:#0f766e,stroke-width:4px,stroke-dasharray:6 3
class SUP,CLS,CLU,DIM,CNN,RNN,TRANS,LLM,GEN,RL,OPT,ENS,NLP,CV,MULTI,RAG,FTUNE,GNN,MLOPS clickable
click SUP "#logistic-regression" "Logistic Regression diagram"
click CLS "#decision-tree" "Decision Tree diagram"
click ENS "#random-forest" "Random Forest diagram"
click CLU "#k-means" "K-Means clustering diagram"
click DIM "#pca" "PCA diagram"
click OPT "#gradient-descent" "Gradient Descent diagram"
click CNN "#cnn" "CNN architecture diagram"
click RNN "#rnn-lstm" "RNN / LSTM diagram"
click TRANS "#transformer" "Transformer attention diagram"
click LLM "#llm" "LLM generation & training diagram"
click GEN "#gan" "GAN & Diffusion diagrams"
click RL "#q-learning" "Q-Learning diagram"
click NLP "#nlp" "NLP pipeline diagram"
click CV "#computer-vision" "Object detection pipeline diagram"
click MULTI "#clip" "CLIP multimodal diagram"
click RAG "#rag" "RAG pipeline diagram"
click FTUNE "#fine-tuning" "Fine-tuning & alignment diagram"
click GNN "#gnn" "Graph neural network diagram"
click MLOPS "#mlops" "MLOps lifecycle diagram"
Scroll to zoom · drag to pan · use the on-map controls to reset. Click a highlighted node to jump to its diagram.
Algorithm Diagrams
How the flagship algorithms actually work — flow & architecture
Logistic Regression
A linear score is squashed by the sigmoid into a probability, then thresholded into a class.
graph LR
X["Input Features $x_1, x_2, \dots, x_n$"]:::io --> Z["Linear Combination $z = \beta_0 + \sum_i \beta_i x_i$"]:::proc
Z --> S["Sigmoid $\sigma(z)=\frac{1}{1+e^{-z}}$"]:::proc
S --> P["Probability $P(y=1\mid x)=\sigma(z)$"]:::proc
P --> T{"$P \geq 0.5$ ?"}:::dec
T -->|Yes| C1["Class 1"]:::io
T -->|No| C0["Class 0"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Decision Tree
Recursive if/else splits chosen to maximize information gain (minimize entropy) until leaves hold a class.
graph TD
R{"Feature A < t₁ ? split on max info gain"}:::dec
R -->|Yes| N1{"Feature B < t₂ ?"}:::dec
R -->|No| N2{"Feature C < t₃ ?"}:::dec
N1 -->|Yes| L1["Leaf: Class 0"]:::io
N1 -->|No| L2["Leaf: Class 1"]:::io
N2 -->|Yes| L3["Leaf: Class 1"]:::io
N2 -->|No| L4["Leaf: Class 0"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Random Forest
Many trees trained on bootstrap samples of the data; their votes are aggregated to reduce variance.
graph TD
D["Training Data"]:::io --> B1["Bootstrap Sample 1"]:::proc
D --> B2["Bootstrap Sample 2"]:::proc
D --> B3["Bootstrap Sample 3"]:::proc
B1 --> T1["Decision Tree 1"]:::proc
B2 --> T2["Decision Tree 2"]:::proc
B3 --> T3["Decision Tree 3"]:::proc
T1 --> V["Aggregate Majority Vote / Average"]:::dec
T2 --> V
T3 --> V
V --> P["Final Prediction"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
K-Means Clustering
Alternates between assigning points to the nearest centroid and recomputing centroids until stable.
graph TD
A["Choose number of clusters $k$"]:::io --> B["Initialize $k$ centroids randomly"]:::proc
B --> C["Assign each point to nearest centroid $\arg\min_j \lVert x_i - \mu_j \rVert^2$"]:::proc
C --> Dn["Recompute centroids $\mu_j = \frac{1}{|C_j|}\sum_{x_i \in C_j} x_i$"]:::proc
Dn --> E{"Centroids moved?"}:::dec
E -->|Yes| C
E -->|No| F["Final Clusters"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Principal Component Analysis (PCA)
Finds orthogonal directions of maximum variance via eigen-decomposition of the covariance matrix, then projects onto the top-k.
graph LR
A["Data Matrix $X$"]:::io --> B["Standardize zero mean, unit variance"]:::proc
B --> C["Covariance Matrix $\Sigma = \frac{1}{n} X^{T} X$"]:::proc
C --> Dn["Eigen Decomposition $\Sigma = U \Lambda U^{T}$"]:::proc
Dn --> E["Sort by eigenvalue keep top-$k$ components"]:::proc
E --> F["Project $Z = X U_k$"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
↑ top
Gradient Descent
The workhorse optimizer: repeatedly step parameters downhill along the negative gradient of the loss.
graph TD
A["Initialize parameters $\theta$"]:::io --> B["Forward pass: compute loss $L(\theta)$"]:::proc
B --> C["Compute gradient $\nabla_\theta L$"]:::proc
C --> Dn["Update $\theta \leftarrow \theta - \eta\, \nabla_\theta L$"]:::proc
Dn --> E{"Converged? $\lVert \nabla L \rVert < \varepsilon$"}:::dec
E -->|No| B
E -->|Yes| F["Optimal $\theta^{*}$"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Convolutional Neural Network (CNN)
Stacked conv + pooling layers extract hierarchical spatial features, then dense layers classify.
graph LR
IN["Input Image $H \times W \times 3$"]:::io --> C1["Conv + ReLU feature maps"]:::conv
C1 --> P1["Max Pool downsample"]:::pool
P1 --> C2["Conv + ReLU deeper features"]:::conv
C2 --> P2["Max Pool"]:::pool
P2 --> FL["Flatten"]:::proc
FL --> FC["Fully Connected"]:::proc
FC --> OUT["Softmax class probabilities"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef conv fill:#fef3c7,stroke:#d97706,stroke-width:2px,color:#000;
classDef pool fill:#ffe4e6,stroke:#e11d48,stroke-width:2px,color:#000;
classDef proc fill:#e0e7ff,stroke:#4f46e5,stroke-width:2px,color:#000;
↑ top
RNN / LSTM (unrolled)
Hidden state carries information across time steps; LSTM gates control what to keep, forget, and output.
graph LR
X1["$x_1$"]:::io --> H1["LSTM Cell forget / input / output gates"]:::proc
X2["$x_2$"]:::io --> H2["LSTM Cell"]:::proc
X3["$x_3$"]:::io --> H3["LSTM Cell"]:::proc
H1 -->|"$h_1, c_1$"| H2
H2 -->|"$h_2, c_2$"| H3
H1 --> O1["$\hat{y}_1$"]:::out
H2 --> O2["$\hat{y}_2$"]:::out
H3 --> O3["$\hat{y}_3$"]:::out
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef out fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Generative Adversarial Network (GAN)
A generator and discriminator play a min-max game: $\min_G \max_D V(D,G)$. G learns to fool D; D learns to spot fakes.
graph LR
Z["Random Noise $z$"]:::io --> G["Generator $G$"]:::proc
G --> FAKE["Fake Sample $G(z)$"]:::proc
REAL["Real Data $x$"]:::io --> D["Discriminator $D$"]:::proc
FAKE --> D
D --> OUT{"Real or Fake?"}:::dec
OUT -->|"backprop loss"| G
OUT -->|"backprop loss"| D
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Diffusion Model
A forward process gradually adds noise to data; a learned reverse process denoises pure noise back into a sample.
graph LR
subgraph FWD["Forward process $q$ — add noise"]
direction LR
X0["$x_0$ data"]:::io --> XM["$x_t$"]:::noise --> XT["$x_T$ pure noise"]:::noise
end
subgraph REV["Reverse process $p_\theta$ — learned denoising"]
direction LR
XT2["$x_T$ pure noise"]:::noise --> XM2["$x_{t-1}$ $p_\theta(x_{t-1}\mid x_t)$"]:::proc --> X0G["$x_0$ generated"]:::io
end
XT -.->|"start sampling"| XT2
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef noise fill:#f1f5f9,stroke:#64748b,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
↑ top
Q-Learning (Reinforcement Learning)
An agent interacts with an environment, receives rewards, and updates a value table toward the Bellman target.
graph LR
A["Agent policy $\pi$"]:::io -->|"action $a_t$"| E["Environment"]:::proc
E -->|"reward $r_t$, next state $s_{t+1}$"| A
A --> U["Update Q-value $Q(s,a) \leftarrow Q(s,a) + \alpha\,[\, r + \gamma \max_{a'} Q(s',a') - Q(s,a)\,]$"]:::dec
U -->|"repeat"| A
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Large Language Model — Generation & Training
Trained in stages (pretrain → supervised fine-tune → RLHF), then generates text autoregressively one token at a time.
graph LR
subgraph TRAIN["Training pipeline"]
direction LR
PT["Pre-training next-token on web-scale text"]:::proc --> SFT["Supervised Fine-Tuning instruction pairs"]:::proc --> RLHF["RLHF / DPO align to human preference"]:::proc
end
subgraph GENr["Autoregressive generation"]
direction LR
P["Prompt tokens"]:::io --> EMB["Embed + Positional"]:::proc
EMB --> BLK["N × Transformer Blocks self-attention + FFN"]:::proc
BLK --> LOG["Next-token logits"]:::proc
LOG --> SM["Softmax + sample (temperature / top-p)"]:::dec
SM --> TOK["Emit token"]:::io
TOK -->|"append, repeat"| EMB
end
RLHF -.->|"deploy weights"| BLK
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
NLP Pipeline
Raw text is tokenized and embedded, passed through a language model, and routed to a task-specific head.
graph LR
T["Raw Text"]:::io --> TK["Tokenize (subword / BPE)"]:::proc
TK --> EM["Embeddings + positional"]:::proc
EM --> MD["Language Model (Transformer encoder/decoder)"]:::proc
MD --> H1["NER head entity F1"]:::out
MD --> H2["Sentiment head polarity"]:::out
MD --> H3["Translation BLEU"]:::out
MD --> H4["Summarization ROUGE"]:::out
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef out fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Computer Vision — Object Detection
A CNN backbone extracts features; a detection head predicts boxes + classes, then Non-Max Suppression removes duplicates.
graph LR
IN["Input Image"]:::io --> BB["CNN Backbone (ResNet / EfficientNet)"]:::proc
BB --> FM["Feature Maps"]:::proc
FM --> AN["Anchors / Region Proposals"]:::proc
AN --> HD["Detection Head class + box regression"]:::proc
HD --> NMS["Non-Max Suppression filter by IoU"]:::dec
NMS --> OUT["Bounding Boxes + Labels"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
CLIP — Contrastive Multimodal Learning
Image and text encoders are trained to align matching pairs in a shared embedding space via a contrastive loss.
graph LR
IMG["Image"]:::io --> IE["Image Encoder (ViT / CNN)"]:::proc
TXT["Text Caption"]:::io --> TE["Text Encoder (Transformer)"]:::proc
IE --> EMB["Shared Embedding Space"]:::proc
TE --> EMB
EMB --> SIM["Cosine Similarity Matrix image ↔ text"]:::proc
SIM --> L["Contrastive Loss pull matched pairs together"]:::dec
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Retrieval-Augmented Generation (RAG)
Documents are embedded into a vector store; at query time the most relevant passages are retrieved and injected into the prompt so the LLM answers from grounded context.
graph LR
DOCS["Documents"]:::io --> CH["Chunk + Embed"]:::proc --> VDB[("Vector DB FAISS / Pinecone")]:::store
Q["User Query"]:::io --> QE["Embed Query"]:::proc
QE --> RET["Retrieve top-k similarity search"]:::proc
VDB --> RET
RET --> RR["Rerank most relevant"]:::proc
RR --> CTX["Build Context query + passages"]:::proc
CTX --> LLM["LLM"]:::proc
LLM --> ANS["Grounded Answer + citations"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef store fill:#ede9fe,stroke:#7c3aed,stroke-width:2px,color:#000;
↑ top
Fine-Tuning & Alignment (LoRA + RLHF/DPO)
PEFT adapts a frozen base model by training tiny low-rank matrices; alignment then steers the model toward human preferences via a reward model (RLHF) or directly (DPO).
graph LR
subgraph PEFT["Parameter-Efficient Fine-Tuning (LoRA)"]
direction LR
BASE["Pretrained LLM frozen weights $W$"]:::proc --> LORA["Train low-rank adapters $\Delta W = B A$"]:::hi --> TUNED["Task-adapted model $W + \Delta W$"]:::io
end
subgraph ALIGN["Preference Alignment"]
direction LR
SFT["SFT model"]:::proc --> RM["Reward Model from human prefs"]:::proc
RM --> RLHF["RLHF (PPO) maximize reward"]:::hi
SFT --> DPO["or DPO / ORPO direct preference opt."]:::hi
RLHF --> AL["Aligned model"]:::io
DPO --> AL
end
TUNED -.->|"start from adapted weights"| SFT
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef hi fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
Graph Neural Network — Message Passing
Each node updates its representation by aggregating messages from its neighbors; stacking L layers grows the receptive field to L hops.
graph LR
G["Graph nodes + edges"]:::io --> AGG["Aggregate neighbors $m_v = \sum_{u \in N(v)} h_u$"]:::proc
AGG --> UPD["Update node $h_v' = \sigma(W\,[\,h_v \,\Vert\, m_v\,])$"]:::proc
UPD --> STK["Stack $L$ layers $L$-hop receptive field"]:::proc
STK --> RO["Readout node / edge / graph task"]:::io
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
↑ top
MLOps Lifecycle
A continuous loop: track training, register versions, optimize for inference, deploy, monitor for drift, and retrain when performance degrades.
graph LR
DATA["Data + Features"]:::io --> TR["Train + Track W&B / MLflow"]:::proc
TR --> REG["Model Registry versioning"]:::proc
REG --> OPT["Optimize quantize / distill / ONNX"]:::proc
OPT --> DEP["Deploy & Serve Triton / vLLM"]:::proc
DEP --> MON["Monitor drift + metrics"]:::dec
MON -->|"retrain trigger"| DATA
classDef io fill:#dcfce7,stroke:#15803d,stroke-width:2px,color:#000;
classDef proc fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#000;
classDef dec fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#000;
↑ top
↑