Split out of #932's closing summary, where this was recorded as post-M2 but never given an issue.
The SKEEP-003 proposal's phases 7 and 8 were always scoped after M2, and the storage model was built so they could land without another rewrite:
- P7 — device placement.
Placement and MemoryDomain exist and AllocationSpec carries them, but nothing chooses. A tensor's home is decided by whoever constructs it, not by a policy that can see the whole model.
- P8 — graph-level planning.
MemoryPlan is arithmetic over a tensor list; it does not know the shape of the computation, so it cannot reuse a buffer whose live range has ended or size the forward slab from actual liveness rather than a worst-case formula.
Why this is a research issue, not a coding one
Both need a decision about where the information comes from. The eager path has no graph to analyse — by construction, it executes as it is called. The recorded path (RecordingExecution → HloGenerator → StableHLO → IREE) does, and IREE already performs its own allocation planning, so P8 may be less "build a planner" than "stop duplicating one".
Worth establishing before designing: what the recorded path already gives us, and whether P7/P8 are wanted for the eager path at all.
Depends on
Nothing blocking. The storage model (#932) landed the types both phases need.
Split out of #932's closing summary, where this was recorded as post-M2 but never given an issue.
The SKEEP-003 proposal's phases 7 and 8 were always scoped after M2, and the storage model was built so they could land without another rewrite:
PlacementandMemoryDomainexist andAllocationSpeccarries them, but nothing chooses. A tensor's home is decided by whoever constructs it, not by a policy that can see the whole model.MemoryPlanis arithmetic over a tensor list; it does not know the shape of the computation, so it cannot reuse a buffer whose live range has ended or size the forward slab from actual liveness rather than a worst-case formula.Why this is a research issue, not a coding one
Both need a decision about where the information comes from. The eager path has no graph to analyse — by construction, it executes as it is called. The recorded path (
RecordingExecution→HloGenerator→ StableHLO → IREE) does, and IREE already performs its own allocation planning, so P8 may be less "build a planner" than "stop duplicating one".Worth establishing before designing: what the recorded path already gives us, and whether P7/P8 are wanted for the eager path at all.
Depends on
Nothing blocking. The storage model (#932) landed the types both phases need.