I managed to compile my port the first example of Nvidia CuTe layout algebra: https://github.com/mratsim/tattletale/pull/46 It's a crazy library with 99.9% template metaprogramming that is used for the highest performance AI inference library (vllm and sglang) for perf-critical code. And you now have access to the core part of it in Nim and it compiles to Cuda, OpenCL, Vulkan, WebGPU.