Two Dialects, One Operation

P-001 P-002 P-008 P-010 P-011

Two Dialects, One Operation

Machine learning vocabulary and psyche vocabulary are not two fields trading metaphors. If reality is computational and learning is prediction at slow timescales, they are two dialects describing one operation, and the correspondence should therefore transfer constraints, not just labels. This page states the identity claim, sets a criterion for telling load-bearing mappings from decorative ones, and pays out the case the framework had been missing: overfitting proper. The regularization coefficient of a learner and the prior weight of a precision-weighted posterior are the same quantity under a monotone reparameterization, which turns the bias/variance axis into the precision-weighting axis and makes the psyche’s two failure poles (memorize everything, generalize nothing versus a prior no evidence can move) predictable rather than merely nameable. A third failure sits off that axis: the self’s report of its own performance is generated by a different process than the one being scored, so the fit cannot be audited from inside. The closing move is structural: overfitting is never a property of a model alone, only of a model against a distribution it has not met yet, which is the same relational shape the framework already gives agency, moral patiency, meaning, and consciousness.


The claim: identity, not analogy

“The mind is like a neural network” is a weak claim, and a cheap one. The framework makes a stronger and more exposed one. Under P-011, computation is not a model of reality but its mode of being; under P-010, learning, adaptation, and perception are one operation distinguished only by the timescale and persistence of the update. Put those together and the ML/psyche correspondence stops being an import. There is one operation, described twice: once in the vocabulary that grew up around gradient descent, once in the vocabulary that grew up around introspection and clinic.

This raises the stakes rather than lowering them. A framework that treats the mapping as metaphor can retreat when a mapping fails (“it was only an analogy”). A framework committed to identity cannot. Every sloppy correspondence becomes a false claim about one operation instead of a strained figure of speech about two. The discipline the identity claim demands is therefore stricter than the discipline analogy demands.

The criterion: constraints, not labels

A mapping is load-bearing when it transfers a constraint: a theorem, a tradeoff, an impossibility result, a prediction that could come out wrong. It is decorative when it transfers only a name.

Load-bearing, from material already in the corpus:

MappingConstraint it transfers
Regularization = prior precisionλ = π_prior / π_data. The bias/variance axis is the precision-weighting axis, so the failure poles are derivable (see below)
TD error, not reward (neuromodulators as H-variables)The signal encodes change in expected value, not hedonic quality. Predicts the dopamine-lesioned rat: starves in front of food, yet eats with evident pleasure when food is placed in its mouth
Compression = generalization (Funes)MDL and the no-free-lunch results make lossless memorization necessarily non-generalizing. Funes is a theorem with a name, not a cautionary tale
One-shot encode plus replay (hippocampus)Complementary learning systems: interleaved replay is required to add fast-learned items without overwriting slow-learned structure. Predicts what breaks when replay is prevented
In-context learningDissolves the train/inference boundary, which is exactly P-010’s content. No separate mechanism for adaptation is needed or permitted
Chain-of-thought as external statePlaces systems on the running/storing continuum (T-003) by what they must buy with tokens when hidden state is absent

Decorative, and worth naming so the identity claim does not become a license:

  • “The brain is a neural network.” Neural is etymology, not homology (see the neural network etymological critique).
  • Transformer attention ≈ biological attention. They share a softmax and an English word. Biological attention is a gain applied to prediction-error channels (precision weighting); transformer attention is content-addressed retrieval over a context window. No constraint crosses.
  • Memory as storage and retrieval. Decorative in both dialects at once: hippocampal encoding plus cortical replay is reconstruction, and a context window is not a database.
  • “Dopamine is the reward chemical.” Not merely decorative but inverted. The load-bearing version is the row above.

The test is cheap to apply and it is the whole hygiene program: ask what the mapping forbids. If the answer is nothing, it is a label.

Regularization is prior precision

The gap the corpus had been carrying: Funes named the memorization pole and P-001 supplied precision weighting, but nothing joined them into a bias/variance account of a psyche. They join tightly, and the join is arithmetic rather than suggestive.

Penalized least squares minimizes ‖y − f(x;θ)‖² + λ‖θ‖². That objective is MAP estimation under a Gaussian likelihood and a zero-mean Gaussian prior over parameters, with

λ = σ²_noise / σ²_prior = π_prior / π_data

The Bayesian brain writes its posterior as a precision-weighted average with

w_prior = π_prior / (π_prior + π_data)

Both are monotone functions of one ratio r = π_prior / π_data: λ = r and w_prior = r/(1+r). The regularization coefficient of a learner and the prior weight of a perceiving brain are the same number, reparameterized. This is a derivation, not an analogy, and it is exactly the kind of transfer the criterion above demands: it forbids things.

Two caveats keep it honest.

Timescale. The identity is cleanest at the learning end, where high-precision priors resist parameter update. At the inference end the same ratio biases the posterior without changing parameters. Under P-010 that is not two phenomena but one ratio applied at two timescales, which is precisely what P-010 asserts, so the framework is not free to use the identity here and disown it there.

λ is not a knob but an estimate. High λ is not pathology; it is what a well-trained model looks like. The pathology is λ that fails to track the reliability of the evidence. So the interesting object is one level up: the expected-precision “shadow hierarchy” that bayesian-brain.md already describes is empirical Bayes performed online, hyperparameter estimation with no outer loop and no held-out set. Precision of precision is where the real inference happens.

The two poles

With r as the axis, the failure modes are the limits.

r → 0 (under-regularized, high variance). Sensory evidence dominates; nothing is compressed; every instance is fitted on its own terms. This is Funes as a computational regime rather than a literary figure: no umwelt forms, because an umwelt is a compression scheme, and no latent variables stabilize, because a latent variable earns its place only by predicting across instances. The predictive-processing literature reaches the same pole from the clinical side: HIPPEA (high, inflexible precision of prediction errors) models autistic perception as prediction error weighted too heavily to be explained away, with the world consequently arriving as detail that will not resolve into kinds.

r → ∞ (over-regularized, high bias). The prior cannot be moved by evidence. This is the canalization side already in the corpus: deep belief attractors, Waddington grooves cut so far down that the marble has no other channel, depression and OCD as priors with too much precision and too little falsifiability.

Reading REBUS through this axis sharpens it into something testable. If psychedelics reduce the precision weighting of high-level priors, they are not delivering content; they are intervening on a schedule, temporarily lowering λ so the system can leave a local optimum. The prediction follows immediately and is falsifiable: therapeutic value should live in what gets relearned inside the low-λ window, not in the experience that window contains. Which is, notably, what the clinical emphasis on integration and on set and setting already behaves as if it believes.

The third failure: fit that cannot be audited

The two poles are failures of fit. There is a third failure, off that axis entirely, and it is a failure of evaluation.

The interpreter emits an account of why the organism did what it did, and that account is not a readout of the computation that produced the behavior. So the self’s report of its own loss is generated by a different process than the one being scored. In ML terms: no held-out set, and the evaluation metric is produced by the system under evaluation. Choice blindness is the receipt. About 80% of subjects fail to notice when their stated preference is swapped, and their justifications for the swapped choice are statistically indistinguishable from genuine ones; in political polling, 92% accept altered responses and 48% become willing to switch coalitions. There is no true-self database being consulted.

The framework already contains the countermeasure, and it is architectural rather than moral. The score is deliberately not writable by the conscious self, “otherwise you would cheat.” Evolution did not solve reward hacking by making the agent honest about its performance. It solved it by withholding write access to the metric. Read that next to the reward-model gaming problem in RLHF and the mapping is load-bearing in the strict sense: it says where to put the metric, not merely what to call the failure.

Overfitting is a relation, not a property

Overfitting cannot be defined against a model alone. It is defined against a distribution the model has not met yet. Two models with identical parameters, identical training loss, and identical everything-you-can-measure-now differ in whether they are overfitted only in virtue of a future they have not encountered.

That is the framework’s characteristic shape, appearing again:

  • Agency is not intrinsic but the residual unpredictability between agent and observer (sphexishness).
  • Moral patiency is a relation to a modeler, not a property of the patient (P-008).
  • Meaning is relational all the way down, with no scaffolding from above and no grounding from below (semantic cosmology).
  • Consciousness is the inside view of a relation among subunits, not a substance any subunit has (P-009).
  • Fundamentality is a property of an observer’s purposes, not of levels of being (no-view-from-nowhere.md).

So the intuition that “overfitting” applies to a psyche was tracking something real, and it was tracking the structure rather than the content: the concept was already relational, which is why it slotted in.

The consequence is not decorative. “Am I overfitted?” is structurally unanswerable from inside at the time of fitting, because answering it requires the test distribution. The felt sense of understanding is therefore not evidence of generalization. It is evidence of low training loss. And low training loss is exactly what both failure poles produce: Funes has none at all, and a canalized prior has none either, because it has quietly stopped sampling anything that could raise it. This is the same epistemic asymmetry the framework meets at the other end, where “is X conscious?” is not answerable from outside. Neither limit is a defect in the method. Both are what it looks like when a property lives in a relation and you try to read it off one term.

What this leaves open

The overparameterized regime. The classical bias/variance U-curve is known to fail where models have far more parameters than samples: past the interpolation threshold, test error can descend a second time (double descent, Belkin et al. 2019). A brain is emphatically in that regime, roughly 10^14 synapses against one lifetime of samples. So the λ axis above is correct as a statement about the prior/evidence ratio and its poles, but the classical picture of an optimum in the middle should not be imported wholesale. What replaces the U-curve for a psyche is open, and it is the most likely place for this page to be wrong.

Catastrophic forgetting. Conspicuously absent from the corpus despite the consolidation machinery sitting right beside it. Connectionist networks trained sequentially overwrite old tasks; humans, mostly, do not. The complementary-learning-systems account (fast hippocampal encoding plus slow interleaved cortical replay) is a candidate answer the framework already holds in pieces. Opened as T-011.

Whether precision weighting is formally the regularizer. The Gaussian case is a derivation. The general case, over hierarchical non-Gaussian generative models with online precision estimation, is not established. Opened as T-010.

  • The Bayesian Brain: supplies the precision-weighting arithmetic this page identifies with regularization, plus the belief-attractor and REBUS material the λ axis reinterprets as a schedule intervention
  • Controlled Hallucination: the phenomenological face of the same machinery; its hallucination spectrum is the r axis viewed from inside experience
  • Intelligence as Self-Modeling: where compression, umwelt, and latent variables are derived from first principles, which is what makes the Funes pole a necessity rather than an anecdote
  • Cephalization from Below: the subbasement machinery (hippocampal one-shot encoding, cortical replay, basal ganglia selection) that the load-bearing mappings and the catastrophic-forgetting question both run through
  • Theory of Mind Is Mind: the interpreter, choice blindness, and sphexishness, which supply the third failure mode and the relational shape of the closing section
  • Computational Being: Claude: the same translation run in the other direction, asking what a frozen-weight transformer inherits and lacks
  • No View from Nowhere: the general form of the closing move, where properties live in relations and there is no term to read them off
  • P-010: Learning is active prediction at long timescales: the prior that licenses treating a learning-rate claim and a perception claim as one claim
  • P-011: Reality is computational: the ontology that makes the correspondence identity rather than metaphor, and therefore holds it to a stricter standard

References

  • Belkin, M., Hsu, D., Ma, S., & Mandal, S. (2019). Reconciling modern machine-learning practice and the classical bias-variance trade-off. PNAS, 116(32), 15849-15854.
  • Borges, J. L. (1942). “Funes el memorioso.” Ficciones.
  • Carhart-Harris, R. L., & Friston, K. (2019). REBUS and the anarchic brain. Pharmacological Reviews, 71(3), 316-344.
  • French, R. M. (1999). Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences, 3(4), 128-135.
  • Johansson, P., Hall, L., Sikström, S., & Olsson, A. (2005). Failure to detect mismatches between intention and outcome in a simple decision task. Science, 310(5745), 116-119.
  • Kirkpatrick, J., et al. (2017). Overcoming catastrophic forgetting in neural networks. PNAS, 114(13), 3521-3526.
  • McClelland, J. L., McNaughton, B. L., & O’Reilly, R. C. (1995). Why there are complementary learning systems in the hippocampus and neocortex. Psychological Review, 102(3), 419-457.
  • McCloskey, M., & Cohen, N. J. (1989). Catastrophic interference in connectionist networks: the sequential learning problem. Psychology of Learning and Motivation, 24, 109-165.
  • Van de Cruys, S., et al. (2014). Precise minds in uncertain worlds: predictive coding in autism. Psychological Review, 121(4), 649-675.