FeedSources
BB

Baseten Blog

19 articles total

Go to source

Stories, updates, and other resources from Baseten.

Go to source
  •  Welcome, Mani Parkhe!
  •  Making Kimi K3 tokenization 18x faster for million-token agentic workloads
  •  How to build a day-0 API for Kimi K3
  •  How we built the new fastest API for GLM-5.2
  •  Introducing GLM 5.2 Fast
  •  How to choose an AI model: lessons from Notion and Gamma
  •  How to optimize LLM inference speed and reduce costs in production
  •  H100 vs. H200 GPUs
  •  GLM 5.2 With Vision
  •  Real-time video generation inference on Baseten
  •  Fast, accurate retrieval with NVIDIA Nemotron 3 Embed
  •  Meet Inkling: Thinking Machines Lab's new customizable model
  •  Introducing Step 3.7 Flash: multimodal reasoning at scale
  •  Building with NVIDIA Nemotron 3 Ultra and LangChain Deep Agents Code on Baseten
  •  H100 vs. H200 vs. B200: which GPU should you use?
  •  How to run GLM-5.2 in any harness
  •  AI training vs. inference: what's the difference?
  •  Live draft model training for speculative decoding
  •  NVIDIA BioNeMo Agent Toolkit on Baseten