In 2008 a group of experienced meditators sat down to a plain task: count your own heartbeats, without taking your pulse. Afterwards they were asked how it had gone. They rated their performance as better than the control group’s, and they found the task easier than the controls found it. Their actual accuracy was no better at all.
That is a real finding and it is worth having. But the more interesting question is the one nobody in the room had to ask, because the answer had already been decided for them: why heartbeats?
Two sentences that sound like one
“I am very attuned to my body” is a self-rating. “I detect my own heartbeat” is a performance. In ordinary speech they pass for the same claim, and in the data they come apart. Where both sides are measured with instruments that are stable over time — retested at .60 and .73 — the correlation between objective accuracy and self-reported body awareness is .06. That is not a weak relationship. That is no relationship.
It is not simply that people flatter themselves, either. Professional dancers do have genuinely elevated accuracy, measurably better than controls, and it still fails to line up with how they describe themselves. The two things are not the same thing.
Nobody in this field ever proposed the heart
Here is what is easy to miss. No somatic tradition has ever claimed that the skill it trains is cardiac enumeration. Not Reich, not Gendlin, not any teacher of any lineage. The heart entered this literature for one reason: it produces a discrete, countable, externally verifiable event several thousand times an hour, and almost nothing else in the interior of the body does.
That is a perfectly good reason to build a paradigm on it. It is not a reason to treat the result as a verdict on body-reading. The criterion was chosen for availability, not because anyone argued it was central — and a conclusion can only travel as far as its criterion reaches.
The measure is in worse shape than its use suggests
Since the whole edifice rests on that one task, its condition matters. In a sample of 572, over 95% of scores were under-counts, and the score turns out to be structurally tied to the participant’s own heart rate. When the instruction is tightened from “count your heartbeats” to “count only the beats you actually feel”, scores roughly halve — meaning a good share of what has been published as interoceptive accuracy was something else. Two published comments argue that part of this critique is an artefact of scoring a ratio variable; both sides accept the under-counting, and the objection about instruction wording has not been answered.
Worse for anyone wanting a clean number: counting your heartbeats and detecting your heartbeats are each internally reliable tasks that do not correlate with each other. There is no single quantity called objective interoceptive accuracy. A systematic review of the field’s own measurement situation calls high validity “potentially elusive”. And the tidy three-part model — accuracy, sensibility, awareness — is not a settled structure either: its only direct test confirmed it in part and found the answer depended on which task was used, and its own senior author moved on in 2022 to a broader framework, on the grounds that assessing those dimensions in isolation does not capture what interoception is.
So why does the dissociation still stand? Because it never depended on that task. It appears where both measures are temporally stable; it appears inside a single sample, in those meditators; and it appears where accuracy genuinely is higher, in the dancers. Demolishing the instrument does not demolish the finding — which is the test of whether a result is real.
What measurement is for is not the same in every discipline
This is our own way of sorting it rather than something taken from the literature, but it does the work that a general complaint about proxies cannot.
In engineering, measurability is the decisive factor — even though the interpretation of proxy quantities is genuinely problematic. Nothing ships on a feeling. The problem is handled by naming it, not by abandoning the measurement.
In science, measurability is also decisive, but a coherent model-based interpretive scheme is more fundamental still. A number without a model that says what the number is a measure of is not a finding; it is a reading. A literature with thousands of measurements and no model bridging its levels is not close to an answer, however precisely it measures.
In phenomenological disciplines, a hard push for measurability on proxy scales can do harm, because it obscures the thing. The available scale decides what will count as the subject, and whatever cannot be displayed on it stops being visible. That is not an argument against measuring. It is an argument about what a measurement here is evidence of.
The failure mode, then, is narrower than “proxies are bad”. It is importing one discipline’s norms into another discipline’s question — and the cardiac literature does exactly that, running a phenomenological question under engineering rules.
And our own side is caught by the same thing
This is where it would be comfortable to say that the felt sense is too rich to be measured, and that the researchers should have come to us. That argument is not available, because it was tried.
The Experiencing Scale has been rating depth of experiencing from therapy transcripts since 1969. It is reliable in good conditions, and it works: a meta-analysis of ten studies and 406 clients finds it predicts outcome, at a magnitude of about r = 0.19 — roughly four per cent of the variance. Something in the neighbourhood of felt sense has been operationalised for sixty years.
And it fails in a structurally identical way. Treatment approach was not a moderator — it predicts outcome just as well inside CBT, and its own authors call it a probable common factor, so it is not measuring anything distinctively experiential. The strongest available critique of its construct validity found its closest correlate to be the density of first-person self-disclosing utterances: a countable feature of speech. Nobody has ever tested whether it is confounded by verbal fluency, education or class — on an instrument that rates transcripts. An independent study in 2026 found it did not change and predicted nothing, while the therapeutic alliance did. And Focusing itself, after sixty years, has no completed randomised trial at all, which is stated in print by the Gendlin Research Center rather than by any critic.
So the honest version is symmetrical, and it is stronger than the self-serving one. Cardiology measures the heart because hearts are countable. The experiential camp measures transcribed speech because speech is recordable. Neither is measuring the construct it names — and whether the absence of a solid model working in simple measurable quantities is a virtue or a failing is a side quarrel about the standing of a discipline, not about its subject.
The same shape, larger
Robert Sapolsky’s Determined (2023) argues there is no free will, on the ground that every decision has neurobiological antecedents all the way down. He does not ignore emergence, chaos or quantum indeterminacy; six chapters address them, and critics who say otherwise are wrong and weaken themselves. The strongest objection to him is not scientific but logical: indeterminism is built into his definition of free will, so the question is settled before the evidence arrives.
David Bentley Hart’s verdict — that the book’s argumentation contains basic logical errors on essentially every page — was delivered in about two minutes of a two-hour conversation with Philip Ball in March 2024, and it is a verdict rather than an analysis. Hart has never written on Sapolsky in print. He is also not a libertarian about free will; he thinks that model is nonsense. His quarrel is with mechanism as metaphysics: with running a question about intention and reason under the rules of a discipline that measures antecedents.
That is the same import error, at a scale where the stakes are the concept of a person.
And the direction the argument must not run
In the 1980s a chain of reasoning looked airtight: ventricular arrhythmias after a heart attack predict sudden death; encainide and flecainide suppress them; therefore they will save lives. The criterion — arrhythmias on a Holter recording — was chosen because it was objective, measurable and quickly available. The phenomenon was survival.
The trial found 56 deaths in 730 patients on the drugs against 22 in 725 on placebo. The drugs did exactly what the surrogate measured, and killed people. The authors of the final report state that the mechanism of the excess mortality remains unknown — which is the part that matters: the proxy’s failure was demonstrated before it was explained. Refusing measurement out of respect for the richness of the phenomenon would not have produced that knowledge. Nothing would have.
So this is not a case for measuring less. It is a case for knowing which question is being asked.
What actually follows
The narrow finding stands and is worth keeping. Confidence in a reading is not evidence of the reading — for the client who says they are very in touch with their body, and equally for the practitioner who is certain about what they are sensing in the room. Those meditators were not lying and were not careless. They were confident, and confidence was uninformative.
Years of practice are not a credential for accuracy. They may be a credential for steadiness, for tolerance of difficult states, for staying present with another person — and those are not nothing. They are simply not the same as reading the body more correctly, and this field lets the first stand in for the second.
What does not follow is that the cardiac result is the verdict on body-reading. It is a verdict on heartbeat counting, which nobody here ever claimed was the skill.
The evidence, including what stands against all of this, is on the What We Know, and How Well page — sources named, gaps marked as gaps.