Isolate whether a model can sum item contributions once their products are already supplied.
Seed 3300 on gpu: 1 of 1 behavioral checks met. See the measurements below. The experiment code has changed since this run; these measurements describe its earlier version.
Insights
The model can sum contributions when their products have already been calculated. This removes value–weight binding and multiplication from the task, leaving numerical encoding, reduction, and decoding. It is useful evidence that the simpler aggregation route works.
A failure on raw pairs alongside success here would point toward the missing interaction, rather than summation alone. However, every collection has the same length, so learning an average and a fixed scaling factor could also solve this task. The proof has an accuracy gate but no separate duplication or permutation gate. Its success should not be described as general arithmetic over arbitrary collection lengths.
Setup
Code
"""A supplied item-local product isolates fixed-width sum reduction.Attention learns a fixed-width sum after products are supplied.Run: uv run python proofs/run.py P015"""from __future__ import annotationsfrom collections.abc import Iteratorimport lightning.pytorch as litimport numpy as npimport torchfrom reporting import reportimport relflow as rfPROOF_ID ="P015"ITEMS =6
Examples
Targets use mask=True; the labels below are supervision hidden from the encoder.
These nonzero contributions cancel. This is another supplied-product diagnostic record, not evidence of a separately tested cancellation guarantee.
Synthetic data and controls
Code
def weighted_records(*, rows: int, length: int, seed: int) -> Iterator[dict]:"""Draw random value/weight pairs and expose raw fields or their products.""" rng = np.random.default_rng(seed)for _ inrange(rows): values = rng.uniform(-1.0, 1.0, size=length) weights = rng.uniform(0.15, 1.5, size=length) contributions = values * weights items = [{"contribution": float(contribution)} for contribution in contributions]yield {"items": items,"weighted_sum": float(contributions.sum()),"weighted_mean": float(contributions.sum() / weights.sum()), }def prediction(model: rf.Model, observations: list[dict]) -> np.ndarray: inputs = [{"items": row["items"]} for row in observations] output = model.predict(inputs)["predictions"].to_pylist()return np.asarray([row["record/weighted_sum"]["content"] for row in output], dtype=np.float64)def rmse(actual: np.ndarray, predicted: np.ndarray |float) ->float:returnfloat(np.sqrt(np.mean(np.square(actual - predicted))))def score(*, train: list[dict], test: list[dict], predicted: np.ndarray) ->dict[str, float]:"""Compare held-out RMSE with the constant training-target mean.""" actual = np.asarray([row["weighted_sum"] for row in test], dtype=np.float64) baseline_rmse = rmse(actual, float(np.asarray([row["weighted_sum"] for row in train], dtype=np.float64).mean())) measured = rmse(actual, predicted)return {"rmse": measured, "baseline_rmse": baseline_rmse, "nrmse": measured / baseline_rmse}
Model tree
Figure 1: Attention reduces six supplied contributions; multiplication is already provided by the inputs.
How it works
The generator computes each contribution = value * weight before the record reaches the model. Learned rf.Attention reductions on the item branch and root combine six contributions into the masked sum target.
This is a diagnostic control for the raw-pair proof. Success establishes fixed-length summation of supplied products. It does not establish that the model learned multiplication or same-item value–weight binding.
Training and evaluation
Code
def run(seed: int, steps: int|None, accelerator: str) ->tuple[dict, dict]: lit.seed_everything(seed, workers=True)# Split seeds are independent; rerunning a generator reproduces the same records. train =list(weighted_records(rows=768, length=ITEMS, seed=seed +1)) test =list(weighted_records(rows=384, length=ITEMS, seed=seed +3)) model = rf.Model( d_model=32, n_layers=2, n_heads=4, reduction=rf.Attention(n_layers=2), batch_size=64, optimizer=lambda module: torch.optim.Adam(module.parameters(), lr=0.001), items=rf.Branch(length=ITEMS, n_layers=2, reduction=rf.Attention(n_layers=2), contribution=rf.Number), weighted_sum=rf.Number(mask=True, objective="mse"), ) datamodule = rf.SyntheticDataModule( model=model, train=lambda: weighted_records(rows=768, length=ITEMS, seed=seed +1), validate=lambda: weighted_records(rows=192, length=ITEMS, seed=seed +2), seed=seed, ) trainer = lit.Trainer( accelerator=accelerator, devices=1, max_epochs=-1, max_steps=500if steps isNoneelsemin(steps, 500), logger=False, enable_progress_bar=False, enable_model_summary=False, enable_checkpointing=False, deterministic=True, num_sanity_val_steps=0, ) trainer.fit(model, datamodule=datamodule)# Evaluate held-out answers and retain their original labels in corruption controls. measured = score(train=train, test=test, predicted=prediction(model, test)) metrics = {"measured": measured, "steps": trainer.global_step} checks = {"Contribution sum nRMSE below 0.25": bool(measured["nrmse"] <0.25)}return metrics, checks
Evidence
Latest full run
Seed 3300, gpu, recorded 2026-09-15T02:22:19.301684+00:00. Outcome: met.
Repeat the control across seeds and test variable lengths and complete-item duplication. Keep the raw-pair proof alongside it so that supplying products does not conceal a missing learned interaction.
The family’s promotion target is at least three core seeds and ten lightweight calibration seeds.
Reproduce
Run by stable ID from the repository root:
uv run python proofs/run.py P015
Or run the self-contained script directly:
PYTHONPATH=proofs uv run python proofs/aggregation/supplied_contribution_sum.py
Add --accelerator gpu for CUDA or --seed 42 for another seeded experiment. --steps 2 checks execution with a short training budget; it is recorded as a smoke run.