Skip to content

Upstream exact Responses prefix routing through llm-d #73

Description

@franciscojavierarceo

Parent tracker: #69
Design and measured evidence: #65

Summary

Deliver the first and highest-impact milestone: exact Responses-aware KV-affine placement through llm-d after agentic-api has rehydrated previous_response_id state.

Scope

  • Upstream vLLM support for accepting ResponsesRequest input at /tokenize and delegating to the same renderer used by inference.
  • Upstream llm-d-router support for dispatching Responses token production to vLLM /tokenize, preserving cache_salt, and rejecting unresolved or incomplete cache identities.
  • Propagate file-discovery rankIndex into endpoint metadata so multiple replicas on one host subscribe to distinct KV-event ports.
  • Configure agentic-api's execution LLM URL to point at llm-d after rehydration rather than routing only on marginal client input.
  • Keep llm-d canonical request keys distinct from vLLM engine block hashes while preserving their event-derived mapping.
  • Verify model, tokenizer, template/parser, salt, LoRA, multimodal identity, block size, and key-format compatibility.
  • Add metrics for matched blocks, selected endpoint, queue/load tradeoff, wrong-pod outcome, and vLLM cached input tokens.

Current evidence

The patched two-replica GPT-OSS-20B proof in #65 produced 12/12 warm returns and 99.08% mean continuation cache hits for both approximate and precise routing, versus 6/12 and 49.54% for load-only. Precise routing reduced mean continuation latency from 811.0 ms to 316.5 ms.

The proof used patched branch images. The vLLM and llm-d-router changes are not yet available in their upstream main branches.

Acceptance criteria

  • Upstream-supported vLLM and llm-d-router images reproduce exact Responses token parity and KV-aware placement without benchmark-only patches.
  • Precise routing materially improves cached tokens and continuation latency over load-only routing under concurrent agent/tool traffic and cache pressure.
  • Different tenant salts cannot share KV blocks; identical permitted salts retain reuse.
  • Wrong-pod selection, event lag, eviction, restart, and AllBlocksCleared are correctness-safe and observable.
  • Same-host and multi-host replica discovery subscribe to the correct KV-event source for every endpoint.
  • Agentic-api does not implement model tokenization or llm-d's canonical hash algorithm.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions