The most expensive domain in medicine is also the most variable.
Musculoskeletal care is the most expensive category of health spending in American medicine — roughly $380 billion a year — and low back pain is both the costliest condition of all and the leading cause of disability in the world. The published evidence shows why the cost runs so high, and what changes when the two things driving it are corrected at the front end, before a problem is treated.
Musculoskeletal care is the largest cost and disability burden in medicine — and the domain with the widest unwarranted variation in how it is delivered.
Musculoskeletal care is the largest category of health spending in the United States — roughly $380 billion a year. Low back and neck pain alone is the most expensive condition in the country, costing more than any single cardiovascular or cancer diagnosis, and low back pain is the single leading cause of disability worldwide. Most of these problems — on the order of nine in ten — are mechanical: they are driven by load and movement, and they are capable of changing with the right movement or management strategy. And most of them enter the system through primary care, which is handed the musculoskeletal problem without being handed a method for solving it.
What makes the domain distinctive is not only its size. It is its variability. How a musculoskeletal problem is assessed, what it is called, and what is done about it vary enormously from one provider to the next — and the cost follows the pathway, not the patient. Two people with the same problem can enter two pathways that differ by an order of magnitude in cost and in outcome, based largely on who they happened to see first. That is the tell. If the variation were coming from the patients, it would track their complexity. It doesn’t. It comes from how they are assessed and managed — which means the problem is in the process, not the person.
The variation has two identifiable sources. Both are fixable — and the rest of this page is about what happens when they are.
The first driver is assessment that is not reliable. Across every discipline that touches musculoskeletal care — physical therapy and chiropractic, primary care, pain management, orthopedics, neurosurgery — the assessment processes in common use have been shown to lack inter-examiner reliability. This is not a criticism of any one profession; the variability lives inside all of them. A handful of individual special tests have respectable sensitivity or specificity, but the assembled examination, taken as a whole, does not produce consistent answers from one clinician to the next — and the problem compounds at scale. The consequence follows directly: assessment variability leads to treatment variability, and treatment variability decreases the chance of a good-to-excellent outcome while raising the cost of care.
This is where the argument returns to first principles. You cannot have validity without reliability. Put plainly: if two trained clinicians examine the same patient and arrive at different conclusions, the examination is telling you more about the examiner than about the patient — and a finding that changes with whoever is holding the patient cannot be a sound basis for treatment. So the precondition for everything else is an assessment that is reliable across clinicians. One has that property demonstrated in the peer-reviewed literature: Mechanical Diagnosis and Therapy. When applied by trained examiners, MDT classification has produced agreement on the order of 95% for the main syndromes and 90% for directional preference, with kappa values reaching 1.0 for lumbar classification — a level of agreement rarely reported anywhere in musculoskeletal assessment research. Reliability is not a feature of the model. It is the foundation the rest of the model stands on.
The second driver is overreliance on imaging. The default use of MRI in musculoskeletal care rests on two assumptions: that what appears on the image is what is causing the pain, and that putting that finding in front of the treating clinician is what tells them how to manage the problem. Patients believe it too — don’t you need an MRI so you know what to do? The published literature has dismantled both. The structural findings used to explain pain — disc bulges, herniations, rotator cuff tears, meniscal tears — are nearly as common in pain-free, fully functioning people as in those who hurt. And imaging a patient while trying to manage the problem conservatively does not help that conservative management; it leaves the patient with a lesser sense of well-being. The cost of the imaging is rarely the scan itself. It is the cascade the scan sets in motion — referral, repeat imaging, injection, procedure, extended disability. This page takes that evidence up in detail below.
These two drivers are the whole of it. Make the front-end assessment reliable, and stop letting imaging substitute for it, and the variation that makes this domain so expensive begins to collapse. What follows is the published evidence for what happens when both are corrected at once — first in low back pain, the hardest and costliest condition there is, then across the rest of the body.
Peer-reviewed claims data from a self-insured employer population, in the condition that costs the most and resists treatment the most: low back pain.
If the two drivers are real, correcting them should show up in the one place that cannot be argued with — what the care costs, and what happens to the patient afterward. It does. And the test case is the hardest one available: not a tidy cohort of fast responders, but unselected low back pain, with and without leg symptoms, in a working population.
Donelson and colleagues, publishing in the Journal of Manual & Manipulative Therapy in 2019, analyzed one year of administrative claims data at a large Global 500 self-insured manufacturer. One group received care guided by a quality-assured mechanical assessment delivered in the employer’s in-house clinic; the other received usual community care. The analysis covered 5,036 patients, none with a low back pain visit in the six months before enrollment, and was risk-adjusted using the payer’s own algorithm. The level of evidence was rated 1b.
After risk adjustment:
Average one-year cost per patient was 51.48% lower in the mechanical-care group.
MRI utilization was 49.75% lower (relative risk 1.99).
Lumbar surgery was 78.38% lower (relative risk 4.73 — a nearly five-fold difference in rate).
Spinal injections were 39.44% lower (relative risk 1.64; trending in the same direction, not statistically significant).
At six months, patients in usual community care were 4.5 times more likely to still be seeking care.
Short-term disability duration was 50% lower — reported separately by the employer, outside the claims analysis.
The surgery figure is the one most easily misread, so it is worth being exact. It is not surgery withheld. It is people who did not need surgery in the first place no longer having it — correctly identified, in advance, as having a problem that resolves under proper conservative management — while the ones who genuinely do need surgery get there sooner, and with better information. The same mechanical-assessment literature shows that among patients already considered surgical candidates, roughly half are found, on a reliable examination, to have a directional preference: a previously undiscovered, mechanically reversible characteristic of their problem. A 78% reduction is not a program rationing surgery. It is a front-end assessment sorting the two groups correctly. Surgery here is not a higher rung the model keeps patients away from; it is the right and necessary step for the patients who genuinely need it.
A claims analysis carries one limitation worth stating plainly: it records what was done and what it cost, not how each clinician reasoned. The data cannot show, visit by visit, how the mechanical-care group managed these patients differently. It does not need to. The cost difference is large, it is risk-adjusted, and the employer had no other cost-saving measure running during the period. When two comparable populations diverge this far in cost and downstream care, and the one systematic difference between them is the assessment-and-management model, the model is the explanation.
The clinicians delivering that care were licensed physical therapists — but the training behind the result is the point. Ninety-six hours of post-graduate MDT education, an additional 350 hours of one-on-one clinical tutoring for many, outcomes-accountable clinician training and examination at a 75% threshold, and enrollment in a data-enabled quality-assurance program. Any case that did not meet its expected response benchmark by the third visit was brought to grand rounds — a structured, collegial review led by an experienced grand rounds tutor, where the group works the case through together: not a flag on the clinician, but the mechanism that keeps every case in front of the problem. That reliable, quality-assured, data-enabled infrastructure is what lets the reliability survive contact with real practice at scale — and it is the same operating system Joint and Spine Solutions runs today.
Joint and Spine Solutions did not conduct this study. The model it uses came directly out of the organization the study evaluated: JSS’s founder was a treating clinician in this work and served as a vice president of clinical development at that integrated mechanical care organization. The methodology, the credentialing, the outcomes accountability, and the grand rounds process the study describes are the same elements that constitute the JSS clinical operating system.
Donelson R, Spratt K, McClellan WS, Gray R, Miller JM, Gatmaitan E. The cost impact of a quality-assured mechanical assessment in primary low back pain care. J Manual & Manipulative Therapy. 2019;27(5):277–286. Level of evidence: 1b.
A substantial share of extremity pain originates in the spine — in patients who have no spinal symptoms and no reason to suspect it. Which makes considering the spine obligatory, not optional.
Donelson establishes what reliable assessment does in low back pain. The same diagnostic logic does not stop at the spine. The most common way a musculoskeletal problem gets mishandled is also the simplest: the pain is treated where it is felt, on the assumption that the location of the symptom is the location of the source. For a meaningful share of patients, that assumption is wrong — and nothing in the presentation warns you, because the person has pain in a shoulder, a hip, or a knee and no pain at all in the spine.
Rosedale and colleagues examined this directly in the EXPOSS study — Extremity Pain of Spinal Source. They took patients presenting with isolated extremity pain who believed their pain was not coming from their spine, and screened each one for a spinal source using a structured mechanical assessment. Across all presentations, 43.5% had a spinal source of symptoms — nearly half of people who were certain the problem was in the limb. Where it matters most clinically, the regional numbers are striking:
Shoulder: just under half — 47.6% — had a spinal source.
Hip: 71%.
Thigh or leg, between the joints: 72%.
Knee: about a quarter — 25.6% — lower, but far from negligible.
The mechanism is somatic referral: a structure in the spine generates pain that is felt out in the limb, with nothing locally wrong at the place that hurts. Because the patient has no back or neck pain, neither the patient nor most clinicians have any reason to look proximally — so the spine is never examined, the extremity is treated on its own, often for a long time, and sometimes a procedure is aimed at a structure that was never the source.
The obligation that follows is modest and exact. It is not that extremity pain is usually spinal — it is not. It is that when someone presents with isolated hip, knee, or thigh pain, or with shoulder pain and a hand on the shoulder blade, doing right by that patient means you have to consider the possibility that the spine is the source, or has some influence on it. At a base rate approaching half for some regions, a clinician who never considers it will miss it in a great many people. Considering the spine is not optional on an extremity presentation. It is obligatory — and it costs one thing: a structured screen. Failing to consider it costs the entire cascade.
That cascade is the same one that makes musculoskeletal care expensive, and the shoulder is the cleanest illustration. Consider the spine, and a shoulder that is in fact cervical can resolve in a few visits with movement directed at the neck. Don’t consider it, and the patient enters the pathway the symptom dictates — examined at the shoulder, medicated, injected, no better, then imaged; the image finds the rotator cuff changes that are common at that age and were never the pain generator, and that incidental finding becomes the target of treatment. It is the two drivers compounding each other: an assessment that never asked the right question, and an image that supplies a convincing wrong answer.
This is why Joint and Spine Solutions screens the spine as a routine part of every extremity presentation — not as a last resort after local treatment fails, but as a question asked at the start.
Rosedale R, Rastogi R, Kidd J, Lynch G, Supp G, Robbins SM. A study exploring the prevalence of Extremity Pain of Spinal Source (EXPOSS). J Manual & Manipulative Therapy. 2019.
Across every region of the body, the structural findings used to explain pain are nearly as common in people who feel nothing. The MRI rarely tells the treating clinician anything that improves conservative care — and it usually leaves the patient with a lesser sense of well-being.
The default use of MRI in musculoskeletal care assumes two things: that the finding on the image is the cause of the pain, and that putting that finding in front of the treating clinician is what tells them how to manage the problem. Patients believe it too — don’t you need an MRI so you know what to do? The strongest evidence on each comes at the question from a different direction, and both assumptions fail.
Start with what knowing does to the patient. Modic and colleagues, publishing in Radiology — the flagship journal of the Radiological Society of North America — ran the study that is hardest to argue with, because it was randomized. Patients with acute low back pain or radiculopathy were assigned either to have their MRI results disclosed within 48 hours or to remain blinded to the findings, with both groups receiving the same conservative care and followed out to two years. The blinded patients did not do worse. They did slightly better. Improvement and satisfaction at six weeks were no different between the groups, and the patients who knew their imaging findings reported a measurably lower sense of well-being. The authors concluded that in typical low back pain, the MRI added no measurable value to planning conservative care, and that knowing the findings did not change the outcome — it only left the patient feeling worse about their condition.
The mechanism is human and predictable. A patient told they have a herniation, a tear, and degeneration now has an explanation for their pain, and they attach the pain to those findings — whether or not any of them is the actual source. The literature does not support that attachment, but the patient cannot unknow it. One detail from the same study makes the point sharper: the patients who had a herniation at the outset actually recovered better than those without one. The finding everyone fears was, if anything, a positive sign.
That is the second assumption — that the finding is the cause — and the asymptomatic literature dismantles it region by region. In people with no pain at all, the very findings blamed for pain — the bulges, the tears, the degeneration — are already there.
Cervical spine.
The neck shows the same dissociation between structure and symptoms. In a study of 1,211 asymptomatic volunteers aged 20 to 70, published in Spine, disc bulging was present in 88% of people with no neck symptoms at all — including roughly three-quarters of those in their twenties — and it grew steadily more common, more severe, and more multi-level with age. The findings that actually warrant concern were uncommon: spinal cord compression appeared in about 5% and cord signal change in about 2%, both rising only after age 50. The everyday finding that gets blamed for neck and arm pain — the bulging disc — is so common in pain-free people that, by itself, it explains nothing.
Nakashima H, Yukawa Y, Suda K, et al. Spine (Phila Pa 1976). 2015;40(6):392–398.
Shoulder.
The same pattern holds, and the strongest evidence is also the most recent. In a population-based study of 602 adults aged 41 to 76, published in JAMA Internal Medicine in 2026, bilateral MRI found a rotator cuff abnormality in 99% of participants — and, shoulder by shoulder, in 96% of the shoulders that had no symptoms at all. Tendinopathy and partial-thickness tears were no more common in painful shoulders than in painless ones. Full-thickness tears were the one finding more frequent in symptomatic shoulders at first glance, but once clinical examination and other imaging findings were accounted for, even that difference disappeared — and 78% of the full-thickness tears found were in shoulders that did not hurt. The authors’ conclusion is blunt: rotator cuff abnormalities are nearly universal after age 40 and correspond so poorly with symptoms that a scan finding, on its own, does not establish what is causing a patient’s pain.
Ibounig T, Järvinen TLN, Raatikainen S, et al. JAMA Intern Med. 2026;186(4):406–414.
Knee.
In healthy volunteers with no history of knee pain or injury, meniscal findings are close to universal. Beattie and colleagues imaged the knees of asymptomatic adults aged 20 to 68 and found that all but one of forty-four had at least one meniscal abnormality, and 61% had abnormalities in at least three of the four regions of the knee — the large majority degenerative changes rather than tears. Among the knees that were normal on X-ray, 97% still showed a meniscal abnormality on MRI. In a much larger population, Englund and colleagues, in the New England Journal of Medicine, found the same truth from the other direction: of the people who turned out to have a meniscal tear on MRI, 61% had had no knee pain, aching, or stiffness in the previous month. The finding does not distinguish the knee that hurts from the knee that doesn’t.
Beattie KA, Boulos P, Pui M, et al. Osteoarthritis Cartilage. 2005;13(3):181–186. — Englund M, Guermazi A, Gale D, et al. N Engl J Med. 2008;359:1108–1115.
This is why imaging is a driver of the variability, not a footnote to it. If a finding appears about as often in the painless as in the painful, the finding cannot tell you, on its own, what is causing the pain. An image ordered before the problem is understood does not locate the pain generator. It produces a list of structural findings — most of them incidental — and the most alarming item on the list becomes the target of treatment. Put plainly: the odds that the thing on the scan is the thing causing the pain are far lower than the report makes it look — to the patient and to the clinician both.
And the cost is rarely the scan. It is the cascade the scan sets in motion. The clearest demonstration is in the spine: early MRI for acute, work-related low back pain has been associated with roughly an eight-fold increase in the odds of subsequent surgery, independent of the severity of the problem. The largest study to date puts hard numbers on the whole cascade. In more than 400,000 primary-care patients with uncomplicated low back pain, those who received an early MRI were, after adjustment, nearly thirteen times as likely to undergo lumbar surgery, more likely to be prescribed opioids, left with slightly worse pain a year later, and ran about 1.4 times the downstream cost. The detail that matters most: this held in a salaried, non-fee-for-service system, where no clinician profited from the scan — which means the cascade is not providers chasing revenue. It is what an early image does on its own, by supplying a structural answer that redirects the entire pathway. Read all of this against the Donelson finding — a 78% reduction in surgery when a reliable, quality-assured, data-enabled assessment comes first — and the two are the same fact from opposite ends. Image first, and surgery becomes far more likely. Put a reliable assessment first, and most of that surgery turns out never to have been needed.
Webster BS, Cifuentes M. J Occup Environ Med. 2010;52:900–907. — Jacobs JC, Jarvik JG, Chou R, et al. Observational study of the downstream consequences of inappropriate MRI of the lumbar spine. J Gen Intern Med. 2020;35(12):3605–3612.
None of this is an argument against imaging. It is an argument about sequence. Imaging, a pain-management consult, a surgical consult — each becomes the right move at a specific moment, and there is a reliable way to know when that moment has arrived. Most of the time, the indication is not a finding on day one; it is the patient who does not get better under conservative care. A reliable, unbiased assessment is what surfaces that — it identifies the patient whose problem does not fit a mechanical pattern, and it does so quickly, before months are lost. Imaging is also essential up front in the narrow set of presentations that raise genuine concern from the start: a suspected sinister cause such as tumor, infection, or fracture, a clear traumatic mechanism, or a progressive neurological deficit. In those situations Joint and Spine Solutions advocates for imaging directly. And surgery is not a higher rung the model keeps patients away from; it is the right and necessary step for the patient who genuinely needs it — reached sooner, and with better information. The problem the evidence describes is narrow and specific: imaging used as a substitute for assessment, ordered before anyone has reliably determined what the problem actually is.
Lumbar spine.
Boden and Jensen established decades ago what has never been overturned: abnormal lumbar MRI findings are common in people who have never had back pain. In Jensen’s asymptomatic subjects, published in the New England Journal of Medicine, a disc bulge was present at one or more levels in 52%, and a protrusion in 27% — in people with no symptoms. Boden found that abnormal scans rose to the majority in subjects over 60. The findings that get blamed for back pain are baseline anatomy in the pain-free population.
Boden SD, Davis DO, Dina TS, et al. J Bone Joint Surg. 1990;72A:403–408. — Jensen MC, Brant-Zawadzki MN, Obuchowski N, et al. N Engl J Med. 1994;331:69–73.
A model built on reliable, unbiased assessment forfeits its credibility the moment it overstates its own evidence. So here is exactly what the published research supports — and where it stops.
Three boundaries are worth stating plainly, because naming them is the same discipline the assessment itself runs on.
The Donelson cost and surgery figures are specific to low back pain. The 51% lower cost, the 78% reduction in surgery, the four-to-five-fold differences in downstream care — those magnitudes were measured in low back pain, the hardest and most expensive condition there is, and the page claims them only there. The case that the approach extends across the body rests on different evidence: that a substantial share of extremity pain is spinal in origin, that the structural findings driving imaging-based care are common in pain-free people in every region, and that the assessment is reliable across joints, not just in the spine. The page makes the cross-region argument on that footing — not by stretching a low-back number to cover a shoulder.
The reliability figures describe trained examiners. The agreement rates that make this assessment dependable were measured in clinicians who had completed the credentialing — they do not transfer to anyone who picks up the method untrained. That is not a hedge; it is the operating model. The reliability is reproducible precisely because it is gated behind training, examination, and ongoing, data-enabled quality assurance. It is a property of a credentialed process, not a technique anyone can claim — which is exactly why the Donelson infrastructure, and the Joint and Spine Solutions model built from it, treats credentialing and grand rounds as non-negotiable rather than optional.
The Donelson study is a claims analysis, not a randomized trial. Its authors are explicit that an analysis of two groups who chose their own care cannot rule out every difference between them, and that a full cost-effectiveness study with baseline clinical data on every patient would be required to attribute the savings to the model with absolute certainty. They make the case — and the evidence supports it — that the model is the most likely explanation, and they claim no more than that. Two things make the design a strength rather than a weakness: claims data captures the real world completely, with no patients lost to follow-up and every cost counted, which a randomized trial rarely achieves; and the result was risk-adjusted, large, and recorded while no other cost-saving measure was running. The page holds the authors’ line exactly — a 51% cost reduction and a four-to-five-fold reduction in surgery and downstream care, at Level 1b evidence, stated at the strength the study earned and no further.
That last point is the whole posture of this page. A program whose entire value is the reliable, unbiased determination of what is actually wrong with a patient cannot then be loose about what is actually true of its own evidence. Claiming precisely what the research supports — and stopping exactly where it stops — is not a limitation on the case. It is the case.
One argument, built in sequence: the most expensive, most variable domain in medicine has two fixable drivers — and the published evidence shows what happens when both are corrected at the front end.
The case this page makes is not a stack of separate studies. It is a single line of reasoning, and each piece of evidence does one job in it.
Musculoskeletal care is the largest category of health spending in American medicine, and low back pain is the leading cause of disability in the world. What makes the domain so costly is not the patients — it is the variation in how they are assessed and managed, variation that does not track their complexity. That variation has two drivers, and the evidence addresses each.
The first driver is unreliable assessment. The reliability literature shows that the response-based mechanical examination, applied by trained examiners, produces agreement across clinicians that the assessments in common use cannot — and that reliability is the precondition for everything downstream, because a process that gives different answers in different hands cannot produce consistent, good-to-excellent outcomes.
The second driver is overreliance on imaging. Across every major region — lumbar, cervical, shoulder, knee — the structural findings used to explain pain are nearly as common in people who feel nothing, so an image taken before the problem is understood does not identify the source; it manufactures a target. And the randomized evidence shows that imaging a patient while trying to manage the problem conservatively does not help that conservative management, and leaves the patient with a lesser sense of well-being.
Then the payoff. When both drivers are corrected at once — a reliable, quality-assured, data-enabled assessment delivered first, which is what tells you whether imaging is even needed — the Donelson claims data show what happens in the hardest and most expensive condition there is. For the patients who received that mechanical care: roughly half the total cost, a near-five-fold reduction in surgery, and half the disability duration. For the patients who went through usual community care instead: four-and-a-half times more likely to still be seeking care six months later. Same condition, same employer, two pathways — and the gap between them is the cost of how the problem was assessed at the start.
The reduction in surgery is not surgery withheld. It is people who did not need surgery in the first place no longer having it — correctly identified in advance as having a problem that resolves under proper conservative management — while the ones who genuinely do need surgery get there sooner, and with better information. And the EXPOSS evidence extends that logic past the low back: a substantial share of extremity pain originates in the spine, in patients with no spinal symptoms at all — which means that considering the spine is not optional on an extremity presentation. It is obligatory.
That is the whole of it. A reliable assessment at the front end, which is what tells you when imaging is actually appropriate; the patients who need a different kind of care identified by the examination itself and confirmed through grand rounds, then moved along quickly; and the cost, the surgery, and the disability that the standard pathway generates simply never accrue.
Joint and Spine Solutions is the application of this evidence into a single coherent model. The clinical philosophy described elsewhere on this site is the same evidence seen from the inside — what it looks like in the room, in how a patient is actually assessed and managed from the very first visit, before anything is treated.
Everything on this page is almost certainly true of your organization right now — and you have probably never been shown it.
Every figure to this point came from someone else’s population — a self-insured manufacturer, a half-million primary-care patients, cohorts of pain-free volunteers. But the pattern they describe is not unique to them. If you carry musculoskeletal risk for a workforce or a patient population, the same two drivers are almost certainly operating in your numbers right now: assessments that vary from one clinician to the next, and imaging ordered before anyone has reliably determined what the problem is. The cost of that is already in your spend. It is simply invisible, because no one has ever measured the quality of your musculoskeletal care — only its total.
That cost is not only imaging and surgery. The same reliance on a structural answer over a reliable assessment runs downstream into medication as well: across the first-contact literature, a physical-therapist-first pathway is associated with lower medication use, and — in every study that has examined it — lower opioid use. Imaging, procedures, prescriptions, disability duration: these are the line items where the cost of an unreliable front end actually lands, and most organizations carry all of them without ever seeing that they trace back to a single point of origin.
This is the question the evidence leaves on your desk. Not whether reliable assessment reduces cost and improves outcomes — the published data settles that. The open question is how much of your current musculoskeletal spend is buying care that was determined reliably at the start, and how much is the cascade that follows when it wasn’t. That number exists. It is in your claims data. Most organizations have simply never been shown it — and being shown it is where this stops being someone else’s evidence and starts being your own.
Cavanaugh AM, Wong M, O’Bright K, Lewis D. The value of integrating physical therapists into primary care, including patient outcomes and health care costs: a scoping review. Physical Therapy (PTJ). 2026;106(6):1–11.
Bring the evidence into your own numbers.
The findings on this page came out of a self-insured employer population whose costs looked like everyone else’s — until the front end changed. If you carry musculoskeletal risk for a workforce, the same question applies to you: what is your current spend actually buying, and how much of it is the cost of how these problems get assessed at the very start?
That is a conversation worth having with the evidence in front of both of us.