lucidrains & zhuzilin were hard working the last days and have completed the following two ring-attention implementations:
Create a test setup that verifies correctness and compares the performance of both solutions.
Phil decided to use a custom triton kernel. Find out why this kernel is used and if it is indeed faster than the cuda flash-attention 2.
Please generate a little report of your findings, either as markdown file or ipynb.