Feed
Sources
Back to sources
BB
Baseten Blog
19 articles total
Go to source
Stories, updates, and other resources from Baseten.
Go to source
Welcome, Mani Parkhe!
Making Kimi K3 tokenization 18x faster for million-token agentic workloads
How to build a day-0 API for Kimi K3
How we built the new fastest API for GLM-5.2
Introducing GLM 5.2 Fast
How to choose an AI model: lessons from Notion and Gamma
How to optimize LLM inference speed and reduce costs in production
H100 vs. H200 GPUs
GLM 5.2 With Vision
Real-time video generation inference on Baseten
Fast, accurate retrieval with NVIDIA Nemotron 3 Embed
Meet Inkling: Thinking Machines Lab's new customizable model
Introducing Step 3.7 Flash: multimodal reasoning at scale
Building with NVIDIA Nemotron 3 Ultra and LangChain Deep Agents Code on Baseten
H100 vs. H200 vs. B200: which GPU should you use?
How to run GLM-5.2 in any harness
AI training vs. inference: what's the difference?
Live draft model training for speculative decoding
NVIDIA BioNeMo Agent Toolkit on Baseten