mstar.engine.resources.attn.flashinfer#
Paged attention through FlashInfer’s prefill/decode wrappers.
Classes
|
- class mstar.engine.resources.attn.flashinfer.FlashInferManager(kv_cache, device, dtype, kv_config, backend='auto')[source]#
Bases:
AttentionManager- property force_double_buffer#
Whether this resource needs double-buffering even off the pre-plan path.
A resource that stages a step’s layout into a reused host buffer and issues a non-blocking H2D into a graph-read device buffer has a race the moment the CPU runs ahead of the GPU: plan(N+1) can overwrite the buffer before step N’s DMA has retired, and the replay attends with N+1’s data. The buffer need not be ours — FlashInfer’s
planholds one per wrapper, which is why the attention managers key theirs per slot. The main runner already double-buffers whenever a resourcesupports_preplan; this flag extends that to a runner that has no pre-plan path but still replays such a resource — notably the piecewise runner — and makes_exec_per_requestfence before it reuses a slot, since its sub-steps all run inside one step. Off by default; a resource opts in only if it has this hazard.
- plan(step, ctx)[source]#
ret is immutable and opaque to runner; only gives to ctx.plan_results
- Parameters:
step (AttentionStep)
ctx (StepContext)
Select last token of the hidden vector per request, used for sampling from prefill.
- property supports_preplan#