flash-attention

Consullo FlashAttention Blackwell fork

This community fork prepares an inference-only CUDA 13 and SM120 FlashAttention wheel for the Consullo Qwen3.8-27B NVFP4 and DFlash2 TensorRT-LLM distribution.