Skip to content

Research 002: test multilayer representational capacity - #5

Open
sfloess wants to merge 5 commits into
mainfrom
research-002-multilayer-capacity
Open

sfloess wants to merge 5 commits into
mainfrom
research-002-multilayer-capacity

Conversation

@sfloess

@sfloess sfloess commented Sep 19, 2026

Copy link
Copy Markdown
Member

Implements issue #4 as a disposable, dependency-free experiment.

  • compares a single-layer perceptron with a two-hidden-unit multilayer learner on XOR
  • uses deterministic initialization and training
  • evaluates fractional held-out examples
  • verifies learned-state restoration
  • avoids Maven, frameworks, Loom, and neural-ai

The experiment is intentionally self-contained. The result should determine whether any abstraction is warranted later.

sfloess commented Sep 19, 2026

Copy link
Copy Markdown
Member Author

Found a concrete acceptance-test inconsistency. The held-out example [0.2, 0.2] -> true is not consistent with XOR: both inputs are on the same side of the binary threshold, so the expected class is false. Running the submitted update algorithm with Java's Random(42L) converges on the four binary XOR examples, but predicts false for [0.2, 0.2], causing the current check(heldOutCorrect == heldOut.size()) to fail. This is a test/data-design problem, not evidence against the learner. Please correct the held-out expected label (or define the fractional held-out labels explicitly from the intended decision regions), rerun the experiment, and record the actual measurements. Keep the experiment disposable and dependency-free as designed.

sfloess commented Sep 19, 2026

Copy link
Copy Markdown
Member Author

Review (Grok)

Reviewed the current PR diff and branch research-002-multilayer-capacity (through 030defa), the experiment markdown, and the surrounding research discipline from Experiment 001. The experiment was compiled and executed; additional instrumentation confirmed the observed behavior.

Focus was experimental validity, not production style.

What works well

  • Disposable, self-contained single Java file. No Maven, frameworks, external ML libraries, Loom, neural-ai, or premature abstractions. This matches the research discipline.
  • Multilayer back-propagation / update equations are correct (tanh hidden, linear output, targets ±1, pre-update output weights used for the hidden delta, standard (1 - h²) derivative).
  • Deterministic initialization, fixed example order, and epoch bound are present and reproducible.
  • Single-layer perceptron correctly fails to represent XOR (plates at 2/4).
  • Multilayer reaches perfect training accuracy on the four binary XOR points under the chosen seed.
  • State capture + restoration into a fresh instance is implemented and checked for prediction equivalence.

REQUIRED (blocks merge / invalidates the claimed evidence)

  1. Held-out accuracy assertion fails for the documented deterministic initialization.

    • With Random(42L) the multilayer fits the training set (4/4) but misclassifies [0.9, 0.9] (predicts true, expected false).
    • The program therefore aborts on check(heldOutCorrect == heldOut.size(), ...).
    • The earlier label correction for [0.2, 0.2] was necessary but insufficient.
    • Seed sensitivity exists: some other seeds (e.g. 53) produce 4/4 held-out as well, but the experiment claims results under the fixed seed that is currently in the source.

    The fractional points encode an informal continuous-XOR expectation that is not entailed by training only on the four vertices with this capacity, loss, and optimizer. Treating perfect held-out accuracy as a hard requirement over-claims what the training evidence shows and currently prevents a clean pass.

  2. The single-layer control is only a partial control for the stated hypothesis (“representational capacity, rather than merely training procedure”).

    • It uses the classic perceptron update + hard threshold, not the identical continuous loss + gradient procedure applied to a linear model.
    • A linear unit trained with the same (output − target) gradient would also be incapable of representing XOR and would make the capacity contrast cleaner. The present control is still informative and matches the documented “perceptron,” but it is not a pure same-procedure ablation.

OPTIONAL

  • Record / emit initial training accuracy (mentioned in the measurements list).
  • Emit the actual per-point held-out predictions so readers can see where the decision surface sits.
  • Document the seed sensitivity of the fractional generalization explicitly.

DEFERRED

Larger capacity, multi-seed success rates, different activations/losses, batch training, or any reusable neural abstractions. These belong in later experiments.

Verdict

The core capacity result on the discrete training set is present and the experiment stays appropriately disposable. However, because the current deterministic run fails its own held-out assertion, the experiment does not produce the passing measurements it claims and is not ready to merge.

Suggested path to merge: adjust the held-out points (or the assertion / documented seed) so that a clean, reproducible pass is obtained under the initialization that appears in the source, keep everything else disposable, and the PR will then constitute valid evidence for the intended research question.

Happy to re-review once the held-out issue is resolved.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant