Floor and Ceil versus Denormals on CPU and GPU

China's fastest gaming GPU still falls far behind RTX 4060

Rotary GPU: Exploring Local Execution for Large MoE Models Under Limited VRAM

China bypasses US GPU bans with 1.54-exaflops 'LineShine' supercomputer — CPU-only monster packs 2.4 million Huawei-designed Armv9 cores

768GB Intel Optane DIMMs to run 1T-parameter LLM with single GPU at 4tps

AMD GPU owners take to Reddit to report fan problem with driver update — Zero RPM feature could cause GPU temperatures to rise unexpectedly

China's new homegrown gaming GPU flops in performance and price — flagship $485 LX 7G100 can't keep pace with Nvidia's older RTX 4060

CubeCL 0.20.0 is Here: Cross-Platform Kernels for Both CPU and GPU

Writing an LLM compiler from scratch [Part 2]: Lowering to a GPU Schedule

GPU Memory Math for LLMs: Formula That Tells You What Fits on Your GPU