CHAPTER 06Disaggregation Flow
Req. prefix: DFv0.4 draft
6.1Pipeline order
DF-01Every entry is disaggregated by four stages in fixed order:
1. Time (entry period → weeks), 2. Attribute→UPC (attribute scope → UPCs, per week),
3. UPC→Product (per UPC-week → product IDs, phase profile or forecast),
4. Location (per product-week → location IDs). Entries already at or below a stage's granularity
pass that stage through unchanged (weight 1 on a single element): a UPC-scoped entry skips stage 2; a
product-scoped entry skips stages 2 and 3.
HI entry (scope, value_su)
│ [1] TIME weekly forecast basis at scope level → value per week
│ [2] ATTR→UPC per week: forecast aggregated per UPC → value per UPC-week
│ [3] UPC→PROD per UPC-week: phase profile | forecast → value per product-week
│ [4] LOCATION per product-week: split table | forecast → value per product-location-week
▼
LO rows (entry_id, product_id, location_id, week_id, quantity_su)
6.2Arithmetic: rounding & exactness
DF-02At every stage, the split of an integer amount over weights MUST use
the largest-remainder method: compute exact shares, floor them, distribute the remaining units one by
one in descending order of fractional remainder.
DF-03Tie-break for equal remainders: ascending lexicographic order of the
element key (week id, UPC id, product id, location id). This makes the pipeline fully deterministic.
DF-04Invariant (per stage and end-to-end): output quantities are
non-negative integers whose sum equals the input amount exactly. Consequently
Σ LO(entry) = value_su for every Disaggregated entry. This is a
mandatory automated test.
DF-05Determinism: identical HI entries, forecast, masters, phase profiles
and split tables MUST produce byte-identical LO row sets, on entry, re-run or batch alike (single code path).
6.3Worked example (end-to-end)
Example E6-A — Entry: attribute brand=B1, ALL locations, no customer, M1, 1 000 SU
Full dataset of Ch. 02, phase profile and split table active. Every number verifiable by hand.
Stage 1 — Time. Weekly basis = all forecast under brand B1 (P1 both locations + P2):
| W1 | W2 | W3 | W4 | Σ |
| basis | 60 | 60 | 90 | 90 | 300 |
| share of 1 000 | 200 | 200 | 300 | 300 | 1000 |
Stage 2 — Attribute→UPC. Per week, forecast per UPC: U1 (=P1) 40/40/60/60, U2 (=P2) 20/20/30/30 — a 2:1 ratio each week:
| Week | amount | U1 exact | U2 exact | → U1 | → U2 |
| W1 | 200 | 133.33 | 66.67 | 133 | 67 |
| W2 | 200 | 133.33 | 66.67 | 133 | 67 |
| W3 | 300 | 200 | 100 | 200 | 100 |
| W4 | 300 | 200 | 100 | 200 | 100 |
W1/W2: floors 133+66=199; the leftover unit goes to the larger remainder (.67) → U2.
Stage 3 — UPC→Product. U1 has no phase profile → forecast fallback → all to P1.
U2 resolves its profile (E5-A):
| Week | U2 amount | rule | → P2 | → P3 |
| W1 | 67 | profile 1.0 / 0 | 67 | 0 |
| W2 | 67 | profile 0.5 / 0.5 | 34 | 33 |
| W3 | 100 | profile 0 / 1.0 | 0 | 100 |
| W4 | 100 | profile 0 / 1.0 | 0 | 100 |
W2: 33.5 each; floors 33+33, tie on remainder .5 → ascending product id gives the unit to P2.
Stage 4 — Location. P1 by forecast mix in W1–W2, by split table 50/50 in W3–W4;
P2 only at L1; P3 by its split-table row set (100% L1). Resulting LO rows:
| Week | P1@L1 | P1@L2 | P2@L1 | P3@L1 | Σ |
| W1 | 33 | 100 | 67 | — | 200 |
| W2 | 67 | 66 | 34 | 33 | 200 |
| W3 | 100 | 100 | — | 100 | 300 |
| W4 | 100 | 100 | — | 100 | 300 |
P1 W1: 133 by mix 10:30 → 33.25/99.75 → remainders send the unit to L2. P1 W2: 133 by 20:20 →
66.5/66.5 → tie → ascending location id → L1 gets 67. Grand total 1 000 = value_su (DF-04). 13 LO rows.
6.4Execution modes
DF-06On-entry: create/update triggers synchronous disaggregation of
that entry alone (OV-02, HE-06). The atomic-replace rule DM-12 applies.
DF-07Batch: re-disaggregates every entry against current inputs.
Triggered by schedule (nightly), after input regeneration (DE-08), or manually from the UI/API/MCP. Entries
are independent (GL-02) — the batch SHOULD parallelise across entries and MUST NOT let one entry's failure
affect others.
DF-08The batch MUST record and expose: start/end time, total wall-clock,
entries processed / succeeded / failed, LO rows written, and per-entry timing distribution (min / median /
p95 / max). This is the demonstrator's centrepiece performance evidence.
DF-09Batch runs are serialised (no two concurrently). Interactive
disaggregation during a batch remains allowed; per-entry atomicity (DM-12) governs, last write wins.
6.5Performance obligations
Reference environment: a small commodity VM — order of 2–4 cores and a few GB of RAM,
single node, local storage. Reference dataset: profile L (100k products, 20 locations, 200 customers,
104-week horizon, 5 000 entries, ≈5 M LO rows; Ch. 10). The exact VM shape used for any
published measurement MUST be stated alongside the numbers.
DF-10Best-effort mandate. HiLo.FM is a technology demonstrator:
on the reference environment it MUST be as fast as the design allows. Fixed performance ceilings are
deliberately NOT specified, so that no implementation choice can be justified by "still within budget" —
an avoidably slower design is non-compliant even if it would meet any figure quoted below.
DF-11Measurement over promises. The demonstrator MUST measure and
expose its actual performance: per-entry timings for interactive disaggregation, the full DF-08 batch report,
and reporting-query timings, all visible in the UI (batch console, Ch. 07) and via API/MCP. Published
figures name the VM shape per §6.5.
DF-12Directional figures (non-contractual). As a direction of
travel on the reference environment: interactive entry disaggregation should feel instant — directionally
p95 ≤ 200 ms for typical entries; a full batch over profile L directionally within ~60 s; LO reporting
queries directionally within ~2 s. These guide design trade-offs and sanity-check measurements; they are
not acceptance thresholds and MUST NOT be cited to justify leaving attainable performance on the table.