FeedSources
AR

AMD ROCm Blog

27 articles total

Go to source

Explore AMD's latest technical deep dives, software optimization guides, and community insights for the ROCm™ platform.

Go to source
  •  AMD GPU Operator v1.5.0: DRA Support, Automated GPU Node Recovery, and Expanded Kubernetes Infrastructure Control
  •  Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
  •  Enabling Language-specific Reasoning in Multilingual Models with Reinforcement Learning
  •  Introducing Instella-MoE: A State-of-the-Art Fully Open Mixture-of-Experts Language Model
  •  Hyperloom - Autonomous Agentic Inference Optimization for AMD GPUs
  •  Introducing AMD ROCm™ Infera: Scaling Goodput for Agentic AI with Distributed Inference Orchestration
  •  Serve Kimi-K2.5-MXFP4 on MI355X with ATOM
  •  Onboard and Deploy Custom Models in AMD AI Workbench
  •  Deploy an Imaging AMD Solution Blueprint on AMD Radeon™ GPUs
  •  Introducing ROCm™ AMD Infinity Context: A Purpose-Built KV Cache Tier for Distributed Inference
  •  Spur: Modern GPU Job Scheduling for HPC and AI Workloads
  •  Scaling MiniMax-M3 Inference with Distributed Serving and Operator Co-Design on AMD Instinct MI355X GPUs
  •  Efficient MiniMax-M3 Inference on AMD Instinct GPUs with ATOM and ATOMesh
  •  Building a High-Performance Video Inference Pipeline with ROCm Libraries Using C/C++
  •  Understanding Attention Algorithms and Their Backends for Image and Video Generation
  •  SPIR-V on ROCm: A Portable IR for AMD GPUs
  •  GEAK V3: Agent-Driven, Repository-Level GPU Kernel Optimization across HIP, Triton, and FlyDSL on AMD GPUs
  •  Multi-Accelerator Support for AIMs and AMD Solution Blueprints
  •  Performance Profiling on AMD GPUs – Part 5: Profiling-Driven Kernel Optimization with an AI Code-Assist Tool
  •  ROCm 7.14: TheRock Goes Production and Expands AMD's AI Software Platform
  •  From Vector Search to Agentic RAG: Building an Enterprise Research Analyst with hipVS
  •  When a Faster Kernel Doesn't Speed Up Serving: Profiling FP8 KV Cache on AMD Instinct MI308X
  •  LogsLop: A Tiny Summarization Tool for Enormous Log Files
  •  Local Image and Video Generation on AMD Ryzen™ AI Max+ Processor (Windows)
  •  GEAK Agent-Driven Optimization of the DeepSeekV4 MLA Kernel
  •  Serving NVFP4 Models on AMD Instinct™ MI355 Accelerators
  •  QuickReduce INT3 Quantization and Benchmarking on MI355
  •  Triton-Based Optimization of Video Sparse Attention on ROCm