Frankenstein Transformer

Overview:

  • Frankenstein Transformer
  • Frankenstein Transformer
  • Frankenstein Transformer
    • Quick Start
    • Feature Matrix
    • Architecture Decision Table
    • CLI Command Reference
    • Mixer Categories
    • Optimizer Categories
    • Documentation Map
    • Installation
      • uv (recommended)
      • pip
      • conda
    • Quick Training Example
    • Vision Transformer (frankenstein_vit) Example
    • License
  • Toolkit Reference
  • Frankenstein Transformer
    • Quick Start
    • Feature Matrix
    • Architecture Decision Table
    • CLI Command Reference
    • Mixer Categories
    • Optimizer Categories
    • Documentation Map
    • Installation
      • uv (recommended)
      • pip
      • conda
    • Quick Training Example
    • Vision Transformer (frankenstein_vit) Example
    • License
  • Transformers Compatibility
  • Transformers export compatibility
    • Integrated --transformers-export flag
    • Export output layout
    • Compatibility matrix
      • High-level training/task compatibility
      • layer_pattern compatibility (from src/schema.yaml enum)

Specifications:

  • Specifications
    • Architecture
    • System Architecture Specification
      • System Design Overview
        • System Architecture Diagram
      • Model Classes
      • Training Modes
      • Looped Depth Formula
      • Layer Pattern Dispatcher
      • Mixture-of-Depths (MoD) Routing
      • Engram Conditional Memory
      • BitNet Path
      • Factorized Embeddings
      • Embedding Convolution
      • Normalization Variants
      • Positional Encodings
      • MoE FFN Routing
    • Attention Mixers
    • Sequence Mixer Families Specification
      • Taxonomy Overview
      • Training-Free Policy
      • Dense Attention Baselines (2)
        • standard_attn — Standard Softmax Attention
        • sigmoid_attn — Sigmoid Self-Attention
      • Grouped-Query Attention (1)
        • gqa_attn — Grouped-Query Attention
      • Recurrent and Retentive Architectures (5)
        • retnet / retnet_attn — Retentive Network
        • mamba — Selective State Space Model
        • ode — ODE-style Continuous Depth Block
        • titan_attn — Titans Memory-Augmented Attention
        • engram_attn — Engram Conditional Memory
      • Sparse Attention Patterns (7)
        • Sparse Attention Design Strategies
        • sparse_transformer_attn — Sparse Transformer
        • longformer_attn — Longformer
        • bigbird_attn — BigBird
        • sparsek_attn — SparseK Attention
        • nsa_attn — Native Sparse Attention (NSA)
        • sparge_attn — SpargeAttn ⚠️ TRAINING-FREE
        • fasa_attn — FASA (Frequency-Aware Sparse Attention) ⚠️ TRAINING-FREE
      • Gated Attention Mechanisms (6)
        • Generic Gating Template
        • gla_attn — Gated Linear Attention
        • deltanet_attn — DeltaNet
        • gated_deltanet_attn — Gated DeltaNet
        • hgrn2_attn — HGRN2
        • fox_attn — Forgetting Transformer (FoX)
        • gated_softmax_attn — Gated Softmax Attention
      • Latent Attention Mechanisms (2)
        • cca_attn — Compressed Convolutional Attention
        • ccgqa_attn — Compressed Convolutional Grouped Query Attention
      • Comprehensive Comparison Table
    • Activation Functions
    • Activation Functions Specification
      • Overview
      • Family Decision Tree
      • Selection Contract
        • Enum (source of truth)
      • ffn_activation_config Keys
      • Rational Activation Function (RAF)
      • GLU Variants
      • Key Files
      • References
    • CLI Reference
    • CLI Command Reference
      • Entrypoint
      • Subcommand Overview
      • Device Choices
      • train — Main Training
        • Examples
      • deploy — Deployment Artifact Creation
        • Examples
      • quantize — Quantized Export
        • Examples
      • infer — Model Inference
        • Examples
      • sbert-train — SBERT Training
        • Examples
      • sbert-infer — SBERT Inference
        • Examples
      • web-server — Streamlit Configuration Builder
        • Examples
      • transformers-export — HuggingFace Export
        • Examples
      • GPU Thermal Guard Flags
    • Deployment
    • Deployment and Quantization Specification
      • Deploy Pipeline
      • Quantization Methods
        • Ternary Weight Packing (BitNet b1.58)
        • Which layers are quantized?
        • Baking faithful ternary weights
        • INT8 Activation Quantization
      • Size Estimates
      • Deployment Artifact Contents
      • Inference Modes
        • Inference Options
      • Quantization CLI
      • Validation
      • HuggingFace Transformers Export
        • BitNet-aware export
      • BitNet.cpp / GGUF Export (best-effort)
        • Scope and limitations
    • mHC: Manifold-Constrained Hyper-Connections
    • mHC: Manifold-Constrained Hyper-Connections
      • Overview
      • Mathematical formulation
      • Implementation (Frankenstein)
      • Config reference
      • Constraints
    • Optimizers
    • Optimizer Families Specification
      • Optimizer Routing Framework
        • Optimizer Selection Decision Tree
      • Prefixed Key Contract
        • Shared Per-Group Suffix Families
        • Optimizer-Specific Global Suffixes
      • Complete Optimizer Inventory
      • Optimizer Family Groupings
      • Key Optimizer Equations
        • AdamW (Baseline)
        • APOLLO Family
        • Anon (Adaptivity-Tunable with IDU)
    • SBERT Workflows
    • SBERT Training and Inference Specification
      • Overview
      • Siamese Training
        • Cosine Similarity Loss
      • Pooling Modes
      • Dataset Types
      • SBERT Configuration Block
      • Inference Modes
        • similarity — Pairwise Scoring
        • search — Top-k Retrieval
        • cluster — Embedding Clustering
        • encode — Embedding Export
      • SBERT Inference Mode Router (Pseudocode)
      • CLI Examples
    • Schema Reference
    • Schema Field Reference
      • Strict Validation Policy
      • Top-Level Sections
      • Model Fields
        • Layer Pattern Valid Values
        • Normalization Valid Values
      • Training Fields
      • Optimizer Fields
        • Valid optimizer_class Values
      • Scheduler Types
      • GPU Metrics Backends
      • SBERT Dataset Types
      • Critical Gotchas
    • Training Safety
    • Training Safety and Stability Specification
      • Gradient Accumulation
      • Global Norm Clipping
      • Post-Clip Explosion Guard
      • NaN/Inf Retry Logic
      • Checkpointing Policy
        • Rolling Checkpoints
        • Best Checkpoints
      • Telemetry
        • CSV Logging
        • GPU Metrics
        • Block Gradient Norms
      • GPU Thermal Guard
        • Thermal State Machine
      • GaLore Controls
      • Scheduler Types
      • Training Step Algorithm (Full)
      • Stability Best Practices
    • Vision Transformer (frankenstein_vit)

API Reference:

  • API Reference
    • CLI Module
      • build_parser()
      • main()
    • Model Package
      • Frankenstein Model
      • Attention Package
        • Common Attention Utilities
        • Standard Attention
        • Sigmoid Attention
        • RetNet Attention
        • ODE Attention
        • Titan Attention
        • Engram Attention
        • Grouped Query Attention
        • Gated Attention
        • Latent Attention
        • Sparse Attention
      • Activation Functions
        • Common Activation Functions
        • Activation Function Factory
        • Exponential Activation Functions
        • GLU Activation Functions
        • Learnable Activation Functions
        • Rectified Activation Functions
      • Embeddings
        • Factorized Embedding
        • HoPE Positional Encoding
        • RoPE Positional Encoding
      • Normalization
        • Normalization Factory
        • DeRF Normalization
        • Dynamic Tanh Normalization
        • RMS Normalization
      • Optimizer Package
        • Optimizer Base
        • Optimizer Factory
        • SGD Momentum
        • AdamW
        • Adafactor
        • GaLore AdamW
        • Prodigy
        • Lion
        • Sophia
        • Muon
        • Turbo-Muon
        • RAdam
        • Adan
        • ADOPT
        • AdEMAMix
        • MARS AdamW
        • Cautious AdamW
        • LAMB
        • Schedule-Free AdamW
        • Shampoo
        • SOAP
        • Anon
        • Apollo
        • Apollo-Mini
        • Q-Apollo
    • Training Package
      • Config Loader
      • Training Main
      • Streaming MLM Dataset
        • StreamingMLMDataset
      • Trainer
    • SBERT Package
      • SBERT Training
      • SBERT Inference
        • SBERTInference
        • SimilarityResult
        • main()
      • SBERT Example Usage
    • Deploy Package
      • Deploy
        • ModelDeployer
        • main()
      • Deploy Inference
        • FrankensteinInference
        • interactive_mode()
        • main()
      • Quantization
        • Design contract
        • ActivationQuantizer
        • BitNetQuantizer
        • bake_bitnet_weights()
        • estimate_model_size()
        • load_quantized_checkpoint()
        • save_quantized_checkpoint()
      • Transformers Export
      • BitNet GGUF Export
      • Deploy Example
    • Tokenizer Package
      • SPM Spa RedPajama35 Tokenizer
        • SpanishSPMTokenizer
    • Streamlit GUI
    • Utils Package
      • Device Utilities
        • is_mps_available()
        • resolve_torch_device()
      • GPU Temperature Guard
        • State machine
        • Backward compatibility
        • GPUTelemetryError
        • GPUTempCheckResult
        • GPUTempSupervisor
        • GPUTemperatureGuard
        • read_temperature_from_nvidia_smi()
      • Schema Loader
      • Storage Manager
        • StorageManager

Bibliography:

  • Bibliography
    • Advanced Sequence Modeling Architectures
    • Advanced Sequence Modeling Architectures: A Comprehensive Theoretical and Empirical Analysis of Transformer Blocks and Emerging Paradigms
      • 1. Standard Attention: The Foundation of Global Contextualization
        • 1.1 Mathematical Formulation and Token Routing Mechanism
        • 1.2 Computational Complexity and the KV Cache Bottleneck
        • 1.3 Architectural Profile: Standard Attention
      • 2. Sigmoid Attention: Hardware-Aware Algorithmic Locality
        • 2.1 Mathematical Foundation and Mixture-of-Experts Perspective
        • 2.2 Stabilization and the Hybrid-Norm Requirement
        • 2.3 Computational Complexity and Hardware Optimization
        • 2.4 Architectural Profile: Sigmoid Attention
      • 3. RetNet: The Duality of Recurrence and Attention
        • 3.1 Mathematical Foundation and the Multi-Scale Retention Mechanism
        • 3.2 Computational Complexity and Performance Scaling
        • 3.3 Architectural Profile: RetNet
      • 4. Mamba: Selective State Space Modeling
        • 4.1 Mathematical Foundation and the Selection Mechanism
        • 4.2 Computational Complexity and the Hardware-Aware Scan
        • 4.3 Architectural Profile: Mamba
      • 5. ODE Transformer: Sequence Generation via Dynamical Systems
        • 5.1 Mathematical Foundation and Runge-Kutta Refinement
        • 5.2 Computational Complexity and the Accuracy Trade-off
        • 5.3 Architectural Profile: ODE Transformer
      • 6. Titans: Test-Time Neural Memorization
        • 6.1 Mathematical Foundation and the Surprise Metric
        • 6.2 Computational Complexity and Infinite Context
        • 6.3 Architectural Profile: Titans
      • 8. Multi-Head Latent Attention (MLA): Low-Rank KV Compression with RoPE
        • 8.1 Mathematical Foundation
        • 8.2 Computational Complexity and the Cache Reduction
        • 8.3 Architectural Profile: MLA
      • 9. Group-Query Latent Attention (GQLA): Hardware-Adaptive Dual Decoding Paths
        • 9.1 Mathematical Foundation
        • 9.2 Computational Complexity and Hardware Adaptivity
        • 9.3 Architectural Profile: GQLA
      • 10. Multi-Head Low-Rank Attention (MLRA): Partitionable Latent for Tensor-Parallel Decoding
        • 10.1 Mathematical Foundation
        • 10.2 Computational Complexity and TP Sharding
        • 10.3 Architectural Profile: MLRA
      • 11. Tucker Attention: A Generalised Low-Rank Factorisation
        • 11.1 Mathematical Foundation
        • 11.2 Computational Complexity and Parameter Efficiency
        • 11.3 Architectural Profile: Tucker Attention
      • 12. Interleaved Head Attention (IHA): Cross-Head Mixing via Pseudo-Heads
        • 12.1 Mathematical Foundation
        • 12.2 Computational Complexity and Reasoning Gains
        • 12.3 Architectural Profile: IHA
      • 13. Grouped-head laTenT Attention (GTA): Shared Maps + Latent Values
        • 13.1 Mathematical Foundation
        • 13.2 Computational Complexity and Cache Compression
        • 13.3 Architectural Profile: GTA
      • 14. Multi-head Temporal Latent Attention (MTLA): Temporal KV Cache Merging
        • 14.1 Mathematical Foundation
        • 14.2 Computational Complexity and Temporal Compression
        • 14.3 Architectural Profile: MTLA
      • 15. Compressed Convolutional Attention (CCA): Attention in a Compressed Latent Space
        • 15.1 Mathematical Formulation and Latent Convolution Tricks
        • 15.2 Computational Complexity and the Compression Factor
        • 15.3 Architectural Profile: CCA
      • 16. Compressed Convolutional Grouped Query Attention (CCGQA): Decoupled Compression with Head Sharing
        • 16.1 Mathematical Formulation and Decoupled Compression
        • 16.2 Computational Complexity and Cache Reduction
        • 16.3 Architectural Profile: CCGQA
      • 17. Synthesis and Systemic Insights: The Future of Sequence Architectures
        • 15.1 The Expressivity versus Compression Duality
        • 15.2 The Dictatorship of Hardware Substrates
        • 15.3 Continuous Dynamics and Geometrical Trajectories
        • 15.4 The Convergence Towards Test-Time Adaptation and Hybridization
    • Sparse Attention Architectures
    • Review of Sparse Attention Blocks in Transformers
      • Executive Summary
      • Sparse Transformer (Child et al., 2019)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • Longformer (Beltagy et al., 2020)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • BigBird (Zaheer et al., 2020)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • FASA — Frequency-Aware Sparse Attention (Wang et al., 2026)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • NSA — Native Sparse Attention (Yuan et al., 2025)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • SparseK Attention (Lou et al., 2024)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • SpargeAttn (Zhang et al., 2025)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • MSA — MiniMax Sparse Attention (Lai et al., 2026)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • SparDA — Sparse Decoupled Attention (Fu et al., 2026)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • Summary Comparison Table
      • References
    • Gated Attention Architectures
    • Gated Attention Blocks in Transformers: A Literature Review
      • Executive Summary
      • 1. Gated Linear Attention (GLA)
        • Description
        • Mathematical Formulation
        • Pros and Cons
        • PyTorch Implementation
      • 2. DeltaNet
        • Description
        • Mathematical Formulation
        • Pros and Cons
      • 3. Gated DeltaNet
        • Description
        • Mathematical Formulation
        • Pros and Cons
      • 4. RetNet (Retentive Network)
        • Description
        • Mathematical Formulation
        • Pros and Cons
      • 5. HGRN2 (Hierarchically Gated Recurrent Network 2)
        • Description
        • Mathematical Formulation
        • Pros and Cons
      • 6. Forgetting Transformer (FoX)
        • Description
        • Mathematical Formulation
        • Pros and Cons
      • 7. Gated Attention (Sigmoid after SDPA)
        • Description
        • Mathematical Formulation
        • Pros and Cons
      • 8. Kimi Delta Attention (KDA)
        • Description
        • Mathematical Formulation
        • Pros and Cons
      • Summary Comparison Table
        • Additional Comparison Dimensions
      • PyTorch Reference Implementations
      • Taxonomy and Evolution
    • Optimizer Families
    • Advanced Optimization Algorithms in Transformer Architectures: A Comprehensive Analysis
      • The Evolution of Optimization in Neural Networks
      • 1. Standard Baseline and Adaptive Optimizers
        • SGD with Momentum
        • Adam and AdamW
        • RAdam (Rectified Adam)
      • 2. Advanced Momentum and Variance Reduction (2024-2025)
        • Adan (Adaptive Nesterov Momentum)
        • ADOPT (Modified Adam with Optimal Convergence)
        • AdEMAMix (Mixture of Two EMAs)
        • MARS (Make Variance Reduction Shine)
        • Cautious Optimizers (C-AdamW, C-Lion)
      • 3. Large-Batch, Memory-Efficient, and Parameter-Free Optimizers
        • LAMB
        • Schedule-Free (AdamW)
      • 4. Second-Order, Geometric, and Orthogonality Optimizers
        • Shampoo
        • SOAP (Shampoo with Adam in Preconditioner’s Eigenbasis)
      • 3. Memory-Efficient and Low-Rank Optimizers
        • Adafactor
        • GaLore (Gradient Low-Rank Projection)
      • 4. Parameter-Free and Distance-Adaptive Methods
        • Prodigy
      • 5. Sign and Second-Order Dynamics
        • Lion (EvoLved Sign Momentum)
        • Sophia
      • 6. Orthogonality-Based Optimizers
        • Muon
        • Turbo-Muon
    • Attention Types Bibliography
    • Optimizers Bibliography
    • Other Bibliography

Technical Reports:

  • Technical Report (English)
  • Informe Tecnico (Espanol)
Frankenstein Transformer
  • API Reference
  • Model Package
  • Activation Functions
  • View page source

Activation Functions

  • Common Activation Functions
  • Activation Function Factory
  • Exponential Activation Functions
  • GLU Activation Functions
  • Learnable Activation Functions
  • Rectified Activation Functions
Previous Next

© Copyright 2026, Erick F. Merino M..

Built with Sphinx using a theme provided by Read the Docs.