RESEARCH & DESIGN NOTES / OCTOBER 2026
The reasoning
behind the story.
Research on the demands of AI inference, the applications driving them, and Tessora’s approach to scaling the system.
[1] COMPUTE AND MEMORY BANDWIDTH
Different serving phases stress different resources
Prompt processing and token generation have different computational patterns. Compute, memory traffic, batch size, precision, and model architecture affect which resource limits performance. We use this distinction to explain why an inference system needs balanced resources.
NVIDIA — Mastering LLM Techniques: Inference Optimization (2023).
[2] MEMORY CAPACITY AND CONCURRENCY
Serving state competes for memory
PagedAttention describes how the KV cache for active requests can consume substantial memory, and how memory management affects batching and serving throughput. The context illustration on our homepage assumes a fixed model, fixed cache precision, independent sessions, and retained context.
The displayed multiplier is (tokens per session × session count) ÷ 16,000. The bars provide a compressed visual scale of proportional KV-cache demand. Model weights and runtime overhead are separate memory requirements; prefix sharing, compression, and attention design affect the cache footprint.
Kwon et al. — Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023).
[3] AGENT WORKLOADS
One user request can expand into many model calls
Anthropic’s research system uses a lead agent and parallel specialists. Its engineering account reports greater token consumption for its agent workflows than for ordinary chats. This informs our agent-fleet scenario and its requirements for concurrent compute, memory, and token throughput.
Anthropic — How we built our multi-agent research system (2025).
[4] LONG-CONTEXT APPLICATIONS
Documents and code become a reusable working set
Google describes context caching for repeated analysis of documents and codebases. Our diligence scenario combines these application patterns with concurrent users to illustrate pressure on memory capacity, memory bandwidth, and prompt processing.
Google Cloud — Save costs and decrease latency while using Gemini with Vertex AI context caching (2025).
[5] SERVING LATENCY AND THROUGHPUT
Useful throughput has a latency requirement
DistServe studies the different demands of prefill and decoding and optimizes serving under response-time objectives. It informs our inference-cloud scenario: operators need both capacity and a serving configuration that delivers acceptable first-token and token-delivery times.
Zhong et al. — DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving (2024).
[6] INTERCONNECT AND SYSTEM DESIGN
Distributed compute introduces communication
DeepSeek’s hardware reflections discuss memory, computational efficiency, interconnection bandwidth, and model–hardware co-design. Together with DistServe, this grounds our discussion of fabric throughput, latency, and communication patterns. Application performance depends on how the compute, memory, interconnect, and serving software work together.
Zhao et al. — Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures (ISCA 2025).
APPLICATION SCENARIO
Private and sovereign AI infrastructure
This is a proposed buyer scenario derived from Tessora’s stated target customers and the resource constraints above. It describes an organization choosing to operate a portfolio of models on infrastructure it controls.
TESSORA SYSTEMS
Architecture and token economics
Tessora’s two core technologies are TSR-Core, its proprietary NPU core technology, and TSR-Link, its direct chip-to-chip fabric. TSR-Core supplies purpose-built inference compute; TSR-Link is designed to connect up to 1,024 chips as one inference system. The architecture combines TSR-Core NPU technology and stacked memory with direct chip connections, a system-wide memory pool, and software partitions for different models and tenants. The TSR-Link fabric uses direct links without internal switches or optical modules. These architecture details come from Tessora’s company profile and selected investor teaser materials.
The 8–10× modeled cost-per-token advantage comes from Tessora’s internal system simulations for selected frontier-scale workloads compared with NVIDIA systems.
The economic model brings together workload and context, token delivery rates, system acquisition cost, utilization, power, facility costs, and communication behavior.
Request a technical briefing