Skip to content

Add sequence-parallel multi-GPU inference (PyTorch and JAX) and a cuDNN flash attention option - #81

Open
weihaokong wants to merge 5 commits into
google-research:mainfrom
weihaokong:pr/pytorch-seqpar
Open

Add sequence-parallel multi-GPU inference (PyTorch and JAX) and a cuDNN flash attention option#81
weihaokong wants to merge 5 commits into
google-research:mainfrom
weihaokong:pr/pytorch-seqpar

Add fused Pallas splash attention option (TPU) for the JAX backend

3c6c587
Select commit
Loading
Failed to load commit list.
Google CLA / cla/google succeeded Jul 23, 2026 in 14s

✅ All contributors are covered under a CLA with Google

See https://cla.developers.google.com/ for more info about Google's Contributor License Agreement (CLA).

ℹ️ Googlers: Go here to view more details and manage scans for this pull request.

Details

The following contributors were found for this pull request:

3c6c587 Author: @weihaokong <kw****o​@gmail.com>

(Only the first commit for a unique contributor is listed.)