flozi00TechHub
Deep dives into servers, GPUs, AI infrastructure and modern IT systems — practical notes from real deployments.
Browse Documentation
Navigate structured guides on hardware, AI infrastructure, Linux and networking.
Find Anything Fast
Full-text search across every guide — specs, configurations and benchmarks included.
in the header — just start typing
Explore by Topic
Guides(54)
- Agent-Fleet Token Economics Is Cache Economics: Decomposing the 60x Dollar Spread
- BOOST and the End of Prefetch: Why Grace Hopper Wants Both Memory Tiers at Once
- The Cache-Read Price War of September 2026 That Wasn't Three-Sided
- CXL-SSDs For LLM Prefix Caching: Byte-Addressability Buys Nothing Without Chunk Awareness
- CXMT at 10% of DRAM Revenue: The DUV Anatomy of a Rise and Its EUV Ceiling
- DeepSeek-V4.1-Flash's KV-Cache Compression, Checked Down to the Byte: 890 Bytes per Token Is a Design, Not a Miracle
- Disaggregated Quantization: Splitting the Prefill and Decode Price System Across Two Formats
- FlashLoop, or: Looping a Transformer Quadruples the Bill and Then Throws Most of It Away
- Gumbel Watermarking Finally Shipped — and the Compliance Asymmetry Nobody Priced
- HBF as the Third KV Tier: 24x Sessions or 5x Worse Latency - the Medium Is Fine, the Placement Policy Decides
- KernelOPT's Agentic Kernel Search: The Verification Cascade Works, The Speedup Fades With Difficulty
- KREX: Shared-GPU Kernel Benchmarking Without Corrupting the Agent's Search
- Who Pays for the KV Cache? Metering Rules Are a Pricing Decision in Disguise
- Risk-Controlled KV Eviction: The Reliability Contract — And the Full-KV Fallback Nobody Ships
- KITE: Scale the Model, Freeze the KV Bill — the Prefill Invariant arXiv:2609.27294 Actually Proves (and the Three Things It Does Not)
- KV-Cache Tensor Decomposition: Where Low-Rank Compression Is Mathematically Dead (and Where It Isn't)
- KVSET: The Oldest Trick in Cache Analysis Just Solved LLM Prefix-Cache Sizing
- Leaderboard Margins vs Hidden Model Selection: Which Gains Survive k Secret Variants?
- Greedy Decoding Is Not Precision-Invariant: BF16-vs-FP16 Flips, Margin Gating, and the Kernel-Order Axis
- Prefix Eviction: Why LRU Is Already Nearly Optimal For LLM Prefix Caching
- PTQ Configuration Pricing: Predicting the Quantization Price Before You Build the Model
- Quantization Breaks Retrieval While Accuracy Holds: The Top-2 Gap Test
- Snapdragon X2 on Linux: The 80-TOPS NPU Is Not the Story — 228 GB/s Is
- Uncheatable Eval: Scoring LLMs by How Few Bits They Need
- Your Agent Fleet Is a Storage Array: The KV-Cache Tiering Math Behind Agentic Inference
- AI Act Article 50 Marking: What the Law Demands on Dec 2, 2026 — and Why Current Watermarking Cannot Meet It Yet
- The Digital Omnibus Didn't Pause the AI Act: What Regulation (EU) 2026/1744 Actually Froze — and Why Mittelstand Is in Inspection Wave One Anyway
- EU AI Act and DSGVO for Self-Hosted LLMs: What Actually Applies
- Why Parakeet TDT Beats Encoder-Decoder ASR
- Random Attention: When KV-Cache Eviction Science Meets a Coin Flip and Loses
- Self-hosting DeepSeek V4 Pro: the actual math
- GPU Buying Guide for LLM Inference (2026)
- Selecting the Right GPU for Qwen3 Inference
- The HBM4 Shortage Is a KV-Cache Economics Problem, Not a Procurement Problem
- How LLMs Actually Work — An Animated Walkthrough
- Germany's KI-MIG Handed the AI Act to the Bundesnetzagentur — Here Is What Your Audit File Needs Before It Asks
- The KV Cache: Bit-Exact Memory Math, GQA vs MQA vs MLA, and PagedAttention
- KV-Cache Glossary: every term in the serving stack
- llama.cpp / Ollama vs vLLM: Which One, When, and Why
- LLM Inference Math: From Theory to Hardware
- LLM VRAM Requirements: A Mathematical Deep Dive
- LoRA and QLoRA Fine-Tuning: The Actual Memory Math
- NVIDIA Groq 3 LPX: The SRAM Decode Engine and the Arithmetic NVIDIA Won't Do
- NVIDIA Buys Hugging Face: The Dependency Math Behind the $12.93B Headline
- LLM Quantization: Bit Layouts, Block Scales, and the VRAM Math
- RAG Retrieval Math: Embedding Memory, Vector Index Size, and the Latency Budget Nobody Calculates
- The 2026 Serving-Engine Churn Tax: vLLM 0.29, SGLang 0.5.19, and TensorRT-LLM Losing TensorRT
- UNISON and the Paused Fleet: Why Agent KV Traffic Breaks Every Cache Policy You Know
- vLLM vs SGLang: Architecture, Overhead, and When to Pick Which
- NVIDIA Driver Issues on Ubuntu 24.04: The nokaslr Fix
- Expert Parallelism: The Router's Dilemma
- Pipeline Parallelism: Pumping Data Upstream
- Tensor Parallelism: Slicing the Silicon
- NVIDIA Drivers, Docker, and GPU Support on Linux
Hardware & Networking(6)
- AMD's Helios 34x Claim, Decomposed: How Much Is Silicon and How Much Is Workload?
- AMD MI300X/MI350/MI400: More Memory, But Does the Math Win? — A Critical Silicon View of AMD Instinct vs NVIDIA
- NVIDIA B200 vs GB200: Efficiency Benchmark
- MLPerf 2026: When the Software Is Worth 2.7x, What Is the Benchmark Measuring?
- NVIDIA NVLink Solutions - High-Speed GPU Interconnect
- Scale-Up vs Scale-Out: The Networking Math Behind LLM Clusters
Latest Articles
How we serve GLM on B300 in production: MTP k=5 tuning
Production notes from serving a GLM-class MoE on 8x B300: the acceptance-rate math behind speculative decoding, why the measured log_stats chain says k=2-k=3 captures most of k=5, and why NVLS failures are silent by default.
Agent-Fleet Token Economics Is Cache Economics: Decomposing the 60x Dollar Spread
The self-reported day of 44 agents, 4 billion tokens and $1,300 versus $80,000 — recomputed line by line against live price sheets, decomposed into cache discount, model spread and looping fraction, and rebuilt as a capacity-planning checklist for anyone running an internal agent fleet.
BOOST and the End of Prefetch: Why Grace Hopper Wants Both Memory Tiers at Once
BOOST (arXiv:2609.13592) decomposed to the bandwidth ledger: why prefetch tiering structurally burns HBM writes during decode, why concurrent proportional access adds the host tier instead of stealing from it, the alpha-math behind +31% throughput and 4.3% TPOT at iso-batch, and the NVLink-C2C scope guard that keeps PCIe x86 hosts out of the claim.
The Cache-Read Price War of September 2026 That Wasn't Three-Sided
The press told a story of three frontier labs cutting prices in unison on September 21-22, 2026. The price sheets say otherwise: one vendor cut cache reads to 0.05x, one shipped new SKUs with an untouched cache multiplier, and one did nothing at all. Decomposed to arithmetic, with per-vendor Barrie-shaped day costings in Python.
CXL-SSDs For LLM Prefix Caching: Byte-Addressability Buys Nothing Without Chunk Awareness
A September 2026 Sogang University paper (arXiv:2609.26828) asks whether CXL-SSDs can replace NVMe for LLM prefix caching and starts with a resounding no: a stock CXL-SSD is about 3× slower than local DRAM and no faster than an NVMe SSD, and a generic next-n prefetcher leaves TTFT essentially unchanged while burning NAND bandwidth. The interesting half is the constructive one — LM-CXD, a co-designed CXL-SSD that makes KV chunks device-visible I/O units, exposes NAND-to-DRAM progress to the serving engine, uses device DRAM as a GPU-accessible buffer, and pipelines layerwise KV movement under GPU compute, reaching up to 2.6× (compute-asynchronous) and 4.03× (layerwise) lower TTFT than the stock device, within 1.5× of local DRAM. This guide prices every step of the interface stack in bandwidth-and-latency arithmetic, works through why CXL's memory semantics do not make loads free, why prefetching without engine knowledge must speculate and lose, and what remains unverified because no CXL-SSD is commercially available yet.