mstar.engine.resources.attn.flashinfer#

Paged attention through FlashInfer’s prefill/decode wrappers.

Classes

FlashInferManager(kv_cache, device, dtype, ...)

class mstar.engine.resources.attn.flashinfer.FlashInferManager(kv_cache, device, dtype, kv_config, backend='auto')[source]#

Bases: AttentionManager

Parameters:
clear_preplan()[source]#
depends_on()[source]#
property force_double_buffer#

Whether this resource needs double-buffering even off the pre-plan path.

A resource that stages a step’s layout into a reused host buffer and issues a non-blocking H2D into a graph-read device buffer has a race the moment the CPU runs ahead of the GPU: plan(N+1) can overwrite the buffer before step N’s DMA has retired, and the replay attends with N+1’s data. The buffer need not be ours — FlashInfer’s plan holds one per wrapper, which is why the attention managers key theirs per slot. The main runner already double-buffers whenever a resource supports_preplan; this flag extends that to a runner that has no pre-plan path but still replays such a resource — notably the piecewise runner — and makes _exec_per_request fence before it reuses a slot, since its sub-steps all run inside one step. Off by default; a resource opts in only if it has this hazard.

plan(step, ctx)[source]#

ret is immutable and opaque to runner; only gives to ctx.plan_results

Parameters:
qo_indptr_buf(label='main')[source]#
Parameters:

label (str)

Return type:

Tensor | None

run(q, label=None, kv_cache_layer=None, k=None, v=None, layer_idx=None)[source]#
Parameters:
Return type:

Tensor

select_last_hidden(hidden, label='main')[source]#

Select last token of the hidden vector per request, used for sampling from prefill.

Parameters:
Return type:

Tensor

property supports_preplan#