The efficient frontier of LLM inference

Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs

A guide to open-source LLM inference and performance

How we got Stable Diffusion XL inference to under 2 seconds

Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA

Show HN: Free Stable Diffusion 2.0 hosted interface

DALL-E Mini – Generate images from a text prompt

Show HN: Baseten – Build ML-powered applications