ARTICLE · INTELLIGENCE

战地情报 · 详情页

来自尧图项目组的一线实战观察与深度解析

MLE 高级 Large-scale matrix operations GPU

MLE 高级 Large-scale matrix operations GPU 结合公开 Systems ML / MLE / GPU-performance 面试经验和真实 AI Infra workload我建议至少能回答下面这些Why can sparse matrix multiplication be slower than dense GEMM?Explain COO vs CSR vs CSC and when you would use each.Why are GNN workloads often memory-bound?How would you optimize neighbor aggregation on a GPU?Explain load imbalance caused by high-degree graph nodes.What is arithmetic intensity?Explain the Roofline model.How do you determine whether a kernel is compute-bound or memory-bound?Why does tiling improve GEMM?What is coalesced memory access?What is a shared-memory bank conflict?What is warp divergence?What is occupancy, and why is 100% occupancy not necessarily optimal?What causes register spilling?Why can kernel fusion improve performance?Why is FlashAttention faster despite computing exact attention?BF16 vs FP16?Why can INT8 improve inference performance?Why doesnt quantization always provide wall-clock speedup?Prefill vs decode bottlenecks?Why is LLM decode often memory-bandwidth-bound?How do you calculate KV-cache memory?MHA vs GQA vs MQA from a systems perspective?Data parallel vs tensor parallel vs pipeline parallel?What collective communication does tensor parallelism require?Why does adding GPUs sometimes reduce efficiency?PCIe vs NVLink matters where?How would you profile a slow inference service?GPU utilization is 30%—how do you investigate?How do you optimize a large matrix operation that does not fit in GPU memory?公开 NVIDIA kernel/performance 相关讨论中memory hierarchy、occupancy、kernel optimization、Tensor Cores、profiling 和 CUDA fundamentals 正是候选人集中准备的主题一则 2026 年相关面试反馈也明确提到 GPU basics 和 kernel-improvement project。
RELATED READING

延伸阅读

更多一线实战笔记与深度复盘,助您持续精进