kernelsJun 30, 2026 · 20 min
Half the tensor cores of a consumer GPU sit idle on dense physics solves. Here's the one line that takes them back.
Dense solves, GP regression, and PDE Green's functions all reduce to one matmul that runs at half rate on a consumer RTX 3080. One line of PyTorch recovers 1.84–1.91× at ~1e-3 error. The reasoning that got there, including where it breaks.
physicstensor coreslinear algebra
1.91×faster dense solvesfp32-acc ▸ fp16-acc