← Back to news section

2026

NeurIPS 2026 Paper Guide: Memorization Is Folding

NeurIPS 2026 Main Track · Poster · CCF Category A

Paper: Memorization Is Folding: Topological Signatures of Noisy-Label Learning.

When training loss keeps falling, is a network learning generalizable structure or memorizing incorrect labels? This paper studies loops in representation space to track how noisy supervision changes a network’s internal structure—and whether those changes can be detected early.

The paper in one minute

  • Question: How does the topology of penultimate-layer representations evolve during noisy-label training?
  • Observation: In some configurations, loop structure fades faster under clean training than under noisy training, producing a crossover between the two trajectories.
  • Use: An early, population-level diagnostic. Estimating noise severity requires calibration for each dataset–architecture pipeline.

01 | Why look beyond loss?

Deep networks can fit noisy labels. A decreasing loss can reflect both learning useful regularities and adapting to incorrect or contradictory supervision. Final accuracy and a loss curve do not directly reveal how internal representations change.

The paper studies penultimate-layer representations. Each sample becomes a point in a high-dimensional feature space; together, these points form a cloud whose connectivity and loops can be followed across training.

02 | Measuring the shape of representations

Persistent homology connects a point cloud at different scales and tracks which structures are short-lived and which persist. H₀ describes connected components; H₁ describes loops. In the paper’s experiments, the key topological discrimination signal under noisy supervision appears in H₁.

The observable is Φ(t) = ∑(d − b), the sum of retained finite H₁ lifetimes. Here b and d denote the scales at which a loop appears and disappears. Bars with lifetime at most 1% of the longest finite lifetime at that checkpoint are removed.

At selected checkpoints, the model is evaluated on test-set inputs and its activations are extracted. The standard setup randomly samples 300 points, computes Euclidean Vietoris–Rips persistent homology with Ripser, and compares clean and noisy Φ trajectories. The paper reports more than 900 training runs and over 8,000 persistent-homology computations.

“Folding” is a geometric interpretation. It concerns changes or persistence of loop structure as the network accommodates contradictory supervision. Φ is not an absolute score that can be compared across arbitrary models without qualification.

03 | A crossover can precede a validation-loss signal

In the figure below, blue denotes clean training and red denotes 50%-noisy training. In crossover-positive configurations, the clean trajectory later declines faster, while the noisy trajectory remains higher for longer. The gap ΔΦ(t) = Φnoisy(t) − Φclean(t) changes from negative to positive.

For CIFAR-10 / ResNet-18, the reported crossover occurs at approximately epoch 25, compared with a validation-loss inflection near epoch 45. These are approximate times because the study samples only 14 checkpoints.

Topological crossover trajectories for clean and 50-percent-noisy training in three configurations

Figure 1 | Representative clean and noisy Φ trajectories. Original Figure 5, PDF p. 7. Open full-resolution figure

Crossover requires more than noisy supervision: the clean signal must fade, the noisy signal must decline more slowly, and the initial gap must close within the training window. Some FashionMNIST runs show the first two trends without crossing within 50 epochs. In the held-out ResNet-34 / CIFAR-10 case, clean training did not enter a fading regime; the predicted absence of crossover was supported in all 9 runs: 3 clean references and 6 noisy runs.

04 | The structure of label contradictions matters

The same noise rate need not produce the same representation dynamics. On CIFAR-10 / ResNet-18, the authors inject 50% noise into either five selected most-confusable or five least-confusable class pairs, using 10 seeds per condition. At epoch 10, the former has 15.9% higher Φ: 19.0±1.5 versus 16.4±0.9, with p=0.0002. The difference weakens later in training.

Topological response to noise in most-confusable versus least-confusable class pairs

Figure 2 | Changing corrupted class pairs while holding the noise rate fixed. Original Figure 3, PDF p. 5. Open full-resolution figure

Noise type also matters. Structured asymmetric confusions that can be absorbed by relatively simple boundary adjustments show weaker folding. Instance-dependent noise and real human annotation noise show stronger signals associated with local contradictions. For the CIFAR-10N worse split, containing roughly 40% noise, the figure shows H₁ increasing from about 12 at epoch 10 to about 27 at epoch 50. Timing depends on architecture: the corresponding noise-type distinction in ViT-Tiny becomes clear only by epoch 50.

Comparison of H1 signals under three realistic noise types on CIFAR-10 and ResNet-18

Figure 3 | Asymmetric, instance-dependent, and CIFAR-10N human noise. Original Figure 10, PDF p. 20. Open full-resolution figure

05 | A complementary signal, with stage-dependent value

The paper compares topology with activation variance, norms, losses, accuracy, and intrinsic dimension. Activation variance and some other statistics are stronger discriminators on average. The value of topology is the additional structural and temporal information it provides.

At early checkpoints, Φ retains information about noise after controlling for the tested statistics. By epoch 50, its unique information relative to training accuracy diminishes, while information relative to test loss remains. The figure illustrates how this complementarity changes during training.

Partial-correlation analysis of topology, test loss, and training accuracy across epochs

Figure 4 | Partial correlations across training stages. Original Figure 2, PDF p. 4. Open full-resolution figure

06 | A reverse check: unfolding during grokking

Grokking is delayed generalization after a model has already fitted its training data. The paper studies modular addition and multiplication with two prime settings and three seeds per setting, giving 12 runs.

Φ decreases during the grokking transition in 10 of 12 runs; the other two are approximately flat. This is consistent with an unfolding interpretation as representations reorganize toward generalizable structure. The evidence here is limited to modular arithmetic.

Ten of twelve modular-arithmetic grokking runs show a decrease in topological persistence

Figure 5 | Changes in Φ across 12 modular-arithmetic grokking runs. Original Figure 8, PDF p. 19. Open full-resolution figure

07 | Early estimation of population-level noise

The paper explores estimating noise severity using Φ(10), measured at epoch 10. After fitting separate calibration relationships for six configurations, it reports R² values of 0.943–0.987, with a mean of approximately 0.962.

The calibration requirements matter: each dataset–architecture pipeline needs a clean reference run and runs at known noise rates. The measurement concerns the population rather than identifying individual mislabeled examples; R² is not a per-sample detection accuracy. On AG News / DistilBERT, the relationship has the opposite sign from the vision settings, so a vision calibration cannot be transferred directly to the text task.

R-squared values for independently calibrated noise-level estimators across six dataset-architecture pipelines

Figure 6 | Fits after separate calibration of six pipelines. Original Figure 11, PDF p. 20; values are R² for population-level noise estimation. Open full-resolution figure

08 | What should readers take away?

This work adds a topological view of training dynamics: alongside loss and accuracy, one can inspect how persistent loops in representations evolve. The results offer directions for dataset auditing, training analysis, and future method design.

  • Crossover is conditional. Its presence depends on trajectory evolution, model and data configuration, and the observation window.
  • Scale and calibration matter. Φ sums Euclidean lifetimes of raw activations, so rescaling changes its absolute value. Different pipelines require separate calibration.
  • The diagnostic is population-level. It does not directly identify mislabeled samples, and lower Φ does not necessarily imply better final accuracy.
  • The central evidence is empirical. The motivating piecewise-linear argument relies on explicit assumptions and is not a universal proof for deep architectures with BatchNorm or residual connections. Full ImageNet and broader settings remain open.

The central takeaway: understanding noisy-label memorization requires reading representation topology together with training stage and the dataset–model configuration.

This guide is based on the supplied manuscript. Figures are reproduced from the paper and identified by their original figure numbers and PDF pages.


Acceptance

The lab’s paper Memorization Is Folding: Topological Signatures of Noisy-Label Learning has been accepted to the NeurIPS 2026 Main Track as a Poster.

The paper’s title focuses on memorization and topological signatures in noisy-label learning.

About NeurIPS

NeurIPS, the Conference on Neural Information Processing Systems, is a major international conference in machine learning and artificial intelligence. Founded in 1987, it is an annual interdisciplinary meeting featuring invited talks, oral and poster presentations of peer-reviewed papers, tutorials, and workshops.

NeurIPS is listed as a Category A international conference in artificial intelligence by the China Computer Federation (CCF). See the official CCF list.

Main Track figures

According to the acceptance notification supplied by the group, the NeurIPS 2026 Main Track received 30,709 valid submissions and accepted 7,900 papers, an acceptance rate of 25.7%. This paper was accepted as a Poster.

The acceptance result and figures above are based on the group’s notification. Venue information is based on the official NeurIPS introduction.