Frankenstein Transformer
Configuration-driven transformer experimentation toolkit with 35+ sequence mixer architectures and 23 optimizer families, strict schema-based configuration, quantized BitNet deployment, and SBERT sentence-embedding training.
Full hosted documentation is at https://frankenstein-transformer.readthedocs.io.
Overview:
Specifications:
- Specifications
- Architecture
- System Architecture Specification
- System Design Overview
- Model Classes
- Training Modes
- Continual Pretraining with
base_model - Looped Depth Formula
- Layer Pattern Dispatcher
- Mixture-of-Depths (MoD) Routing
- Engram Conditional Memory
- BitNet Path
- Factorized Embeddings
- Embedding Convolution
- Normalization Variants
- Positional Encodings
- MoE FFN Routing
- Attention Mixers
- Sequence Mixer Families Specification
- Taxonomy Overview
- Training-Free Policy
- Dense Attention Baselines (2)
- Grouped-Query Attention (1)
- Recurrent and Retentive Architectures (6)
- Sparse Attention Patterns (9)
- Gated Attention Mechanisms (8)
- Fast-Weight Attention / Online Continual Learning (6)
- Latent Attention Mechanisms (10)
- Geometric Field Attention (1)
- Comprehensive Comparison Table
- Config knobs and example assembly
- Activation Functions
- Activation Functions Specification
- CLI Reference
- CLI Command Reference
- Deployment
- Deployment and Quantization Specification
- mHC: Manifold-Constrained Hyper-Connections
- mHC: Manifold-Constrained Hyper-Connections
- Optimizers
- Optimizer Families Specification
- SBERT Workflows
- SBERT Training and Inference Specification
- Schema Reference
- Schema Field Reference
- Training Safety
- Training Safety and Stability Specification
- Vision Transformer (frankenstein_vit)
API Reference:
Bibliography:
- Bibliography
- Advanced Sequence Modeling Architectures
- Advanced Sequence Modeling Architectures: A Comprehensive Theoretical and Empirical Analysis of Transformer Blocks and Emerging Paradigms
- Sparse Attention Architectures
- Review of Sparse Attention Blocks in Transformers
- Gated Attention Architectures
- Gated Attention Blocks in Transformers: A Literature Review
- Optimizer Families
- Advanced Optimization Algorithms in Transformer Architectures: A Comprehensive Analysis
- Attention Types Bibliography
- Optimizers Bibliography
- Other Bibliography
Technical Reports: