←Back to NewsAI News/ImagemodelImageGpusNVIDIA Ships Qwen2.5-VL Optimized for 3.6x Smaller Blackwell GPU InferenceNVIDIA published a 4-bit NVFP4 build of Qwen2.5-VL-7B-Instruct that runs on Blackwell Tensor Cores through TensorRT-LLM, cutting memory roughly 3.5x.SourceNVIDIA AIPublishedSep 29, 2026, 4:02 PMAuthorAlphaSignal NewsroomRead1 min readNVIDIA published a 4-bit NVFP4 build of Qwen2.5-VL-7B-Instruct that runs on Blackwell Tensor Cores through TensorRT-LLM, cutting memory roughly 3.5x.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsAlphaSignal · modelCommunity Strips Qwen-Image 2.1's Refusals to Run Locally on 16GB LaptopsAlphaSignal · repoSolo Dev Runs NVIDIA's DLSS 5 Neural Rendering on AMD Radeon GPUsAlphaSignal · modelJev-Omni Shrinks a 50GB AI Decision Model to Run on Laptop GPUs