Guides & field notes

Writing about AI you own

Practical guidance on running AI on your own hardware: what to buy, what it costs to power, how fast it really is. Every number here carries its denominator, and where something was not measured, we say so.

Decision guide

When to buy an AI cluster, and what to actually expect

Written for the person signing the cheque. What tokens per second really means, how fast Claude and ChatGPT actually are, a real GPU-versus-power-versus-throughput matrix, and which of five tiers fits chat, batch APIs or coding agents.

27 Jul 2026 · 18 min read

Field notes

500K tokens of context, on four boxes you can fit on a desk

A measured deployment of GLM-5.2 across four DGX Sparks: adaptive speculative depth, full CUDA graph coverage, and the OOM that set the memory envelope.

27 Jul 2026 · 9 min read

Guide

On-premise AI server: the enterprise guide

What an on-premise AI server is, how it works, what hardware it needs, and what it takes to run private LLMs on hardware you own.

23 Jul 2026 · 7 min read

How-to

How to run open-source LLMs on-premise

Choosing models, sizing GPU memory, picking an inference engine, and serving a private OpenAI-compatible API on hardware you control.

23 Jul 2026 · 6 min read

Comparison

On-premise vs cloud AI: cost, security & compliance

Total cost over five years, data security and compliance compared, and the honest answer to when self-hosting actually beats a cloud API.

23 Jul 2026 · 6 min read

Ready to size your own stack?

Configure your stack Talk to us