Skip to content
#

torch-neuronx

Here is 1 public repository matching this topic...

Production LLM pipeline on AWS Trainium and Inferentia: LoRA fine-tune Llama 3.1 8B on a trn1.2xlarge, ship the adapter through S3, serve it with vLLM on an inf2.xlarge, and measure everything (TTFT/TPOT percentiles, tokens/s, MFU, goodput at SLO) with compile costs included and failures recorded as receipts.

  • Updated Aug 6, 2026
  • Python

Improve this page

Add a description, image, and links to the torch-neuronx topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the torch-neuronx topic, visit your repo's landing page and select "manage topics."

Learn more