<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Road to CEPAS 2026]]></title><description><![CDATA[Mijn persoonlijke Substack]]></description><link>https://road2cepas.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Bj74!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a8b5807-c726-4759-b248-e1fc37c301a0_1175x783.jpeg</url><title>Road to CEPAS 2026</title><link>https://road2cepas.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 15 Aug 2026 15:06:13 GMT</lastBuildDate><atom:link href="https://road2cepas.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Quantum Neonatology]]></copyright><language><![CDATA[nl]]></language><webMaster><![CDATA[road2cepas@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[road2cepas@substack.com]]></itunes:email><itunes:name><![CDATA[Quantum Neonatology]]></itunes:name></itunes:owner><itunes:author><![CDATA[Quantum Neonatology]]></itunes:author><googleplay:owner><![CDATA[road2cepas@substack.com]]></googleplay:owner><googleplay:email><![CDATA[road2cepas@substack.com]]></googleplay:email><googleplay:author><![CDATA[Quantum Neonatology]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[From Alarm to Better Outcomes]]></title><description><![CDATA[Road to CEPAS 2026, Episode 6]]></description><link>https://road2cepas.substack.com/p/from-alarm-to-better-outcomes</link><guid isPermaLink="false">https://road2cepas.substack.com/p/from-alarm-to-better-outcomes</guid><dc:creator><![CDATA[Quantum Neonatology]]></dc:creator><pubDate>Mon, 27 Jul 2026 17:01:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hPwv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2cedfab0-e3e4-4e90-a33c-a1a27a902cf8_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;images&quot;:[{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2cedfab0-e3e4-4e90-a33c-a1a27a902cf8_1536x1024.png&quot;}],&quot;caption&quot;:&quot;&quot;,&quot;alt&quot;:&quot;Graphic overview &quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2cedfab0-e3e4-4e90-a33c-a1a27a902cf8_1536x1024.png&quot;}},&quot;isEditorNode&quot;:true}"></div><p></p><p>Episode 5 ended with a statement. Our paper showed that the model can separate higher-risk periods from lower-risk with useful but far from perfect accuracy, it flagged roughly half of late-onset sepsis episodes before clinical suspicion without instant alarm fatigue. That is the discrimination question, the easier half. The harder half, does acting on that alarm at three in the morning leave the baby better off? This is still open, because the paper does not answer it and could not have. That is not a modeling question, it is a trial question. This episode is that quest, to go to somewhere concrete.</p><p>The place it leads is a single study run between 2004 and 2010, and the uncomfortable fact that, fifteen years on, it is still pretty much the only one of its kind.</p><h2>The rung we are standing on</h2><p>In the discussion of our 2023 model, we sketched the road a prediction like ours has to travel to earn a place at the bedside. It runs in three stages: retrospective development and internal validation. (which is where our published model sits) Then prospective silent validation, confirming the score stays accurate on new patients while clinicians cannot yet see it; and only then a clinical-impact trial, which asks the different question of whether displaying the score, or acting on it, <strong>actually changes what happens to the baby</strong>. An AI model is a retrospective simulation. It looks back over years of saved oxygen-saturation and heart-rate data, replays them minute by minute, and asks how often the alarm would have fired before the clinician noticed anything. The answer we know, but a replay is not a trial. Care has not changed. No baby was treated differently because the number went up.</p><p>Where this episode stands is the next step, the clinical-impact trial, the patient outcomes, the very thing a good model cannot deliver on its own, and it is where almost nobody in this field has climbed. To understand what it costs come there, it helps to look at the one team that did.</p><h2>HeRO(&#8216;s) trial</h2><p>Between April 2004 and May 2010, across nine American NICUs, 3,003 very low birth weight infants were randomised to have a continuously updated heart-rate risk score either displayed at the bedside or recorded but hidden (<a href="https://doi.org/10.1016/j.jpeds.2011.06.044">Moorman et al., 2011</a>). The score was the HeRO index: a number expressing the fold-increase in the risk of sepsis over the next 24 hours, derived from subtle changes in heart rate, reduced variability and transient decelerations, that tend to appear before a preterm infant looks clinically septic. It is not our model. But it is the same species of thing: a continuous physiological risk score, updated hourly, sitting on a monitor next to a sick baby.</p><p>The result that everyone remembers is a mortality reduction. Infants whose score was displayed died less often, 8.1% versus 10.2%, a hazard ratio of 0.78 and an estimated number-needed-to-monitor of 48. In the pre-specified subgroup of infants below 1,000 grams, the estimated benefit was larger, falling from 17.6% to 13.2%, a number-needed-to-monitor of 23. For a device that changes nothing about the baby except what the clinician can see, that is a striking finding, and it is the reason HeRO is the reference point for this entire conversation.</p><h2>What the trial actually measured</h2><p>There is a wrinkle worth being precise about and it matters more here than anywhere.</p><p>Mortality was not the trial&#8217;s primary outcome. It was a secondary one. The registered primary endpoint was a composite, the number of days a baby was alive and off the ventilator in the 120 days after randomisation, and on that endpoint the trial did not reach significance. Displayed-score infants gained an average of 2.3 such days, with a p-value of 0.08. By the trial&#8217;s own pre-specified terms, that is a miss.</p><p>So the canonical trial in this field is, read strictly, both positive and negative at once: it missed on the outcome it was built to measure and hit on one listed beneath. This holds two possibilities open at the same time. A composite can obscure a real mortality effect if its components move in different directions, the ventilator half adding noise that drags the whole below the threshold. Equally, a significant secondary finding after a negative primary endpoint can itself be a chance result, and mortality was one of several secondaries. The confidence interval for the mortality benefit only just excluded no effect; the p-value was 0.04. A Dutch systematic review weighing exactly this rated the trial&#8217;s risk of bias high, on performance and detection bias and the absence of correction for multiple testing, and concluded that the evidence does not justify putting HRC monitoring into routine care without another trial (<a href="https://doi.org/10.1159/000531118">Koppens et al., 2023</a>). None of this makes the finding worthless. It makes it what it is: clinically important, internally coherent, and statistically less secure than its prominence in the later literature suggests.</p><p>The point that survives either reading is the one that matters most for designing the next study: how much rides on a decision made before a single baby is enrolled, which endpoint you name as primary. The HeRO team reasoned their way to a composite that seemed clinically sensible in 2004. Whether that was the right call is still arguable, which is precisely why the endpoint is the first thing a validation study has to get right, and the thing hardest to know in advance.</p><h2>The counterfactual problem, which is the real difficulty</h2><p></p><p>The HeRO trial deliberately did not tell anyone what to do. Clinicians were taught what the score meant and told that a rising number should prompt them to look at the baby and consider tests, but no specific action was required by the protocol. The authors call this, in their own discussion, &#8220;a debatable weakness.&#8221; I would put it more strongly: it is the exact gap this whole series is about. The trial tested the display of a number. It did not test a defined response to that number. If the mortality effect was causal at all, it must have operated through nine units&#8217; worth of clinicians each responding, in their own way, to the score in front of them.</p><p>And you cannot cleanly measure what that response was worth. To know whether the alarm helped, you would want to know how much earlier sepsis was caught, but the biological moment a preterm infant&#8217;s sepsis begins is unobservable; there is no clock that starts. So time-to-onset, the thing you would most want to anchor to, cannot be measured directly, and the trial says as much: the mechanism of the benefit could not be proven, only inferred.</p><p>That is not a counsel of despair. A new trial cannot see the true onset of sepsis, but it can measure the operational chain the alert is meant to shorten: time from alert to bedside assessment, from alert to blood culture, from alert to antibiotics; physiological deterioration in the hours after an alert; escalation to organ support; adherence to whatever response protocol is specified. None of that pins down biological onset, but it would make the mechanism far less opaque than it was in 2011.</p><p>What the 2011 team could show is indirect, and it is worth laying out because it is the texture a real impact trial lives in:</p><ul><li><p>The amount of proven sepsis was the same in both arms. Displaying the score did not catch <em>more</em> sepsis: 358 affected infants in the displayed group, 379 in the control group, no meaningful difference.</p></li><li><p>Among the babies who did develop proven sepsis, mortality in the following 30 days was lower where the score had been displayed, 10.0% versus 16.1%. The score did not find more infections; the most that can be said is that it may have helped clinicians act sooner on the infections already there, and &#8220;helped&#8221; is an interpretation the design cannot confirm.</p></li><li><p>That came at a cost, though a less certain one than the mortality figures. Displayed-score infants underwent slightly more blood cultures, a difference of borderline significance, and accumulated slightly more antibiotic days, a difference that was not statistically significant at all.</p></li></ul><p>&#8220;Act sooner on the infections already there&#8221; is as far as the strongest trial in the field can take the claim. That is not a failure of the trial. It is the limit of what a trial of this design can prove, and it is the same limit any future study, including any study of our model, would run into.</p><h2>So what would a real trial of our model take</h2><p>Lay the HeRO experience against our 2023 paper and the shape of the required study comes into focus. It is sobering.</p><ul><li><p><strong>A response protocol, not just an alarm.</strong> Our published alarm framework, an eight-hour refractory period and escalating thresholds, defines when the number speaks, not what the clinician does when it does. If the goal is a <em>validated response</em>, and not simply another displayed number, that response has to be written, agreed and fixed before enrolment, so that the thing under test is the whole pathway from alert to action.</p></li><li><p><strong>An endpoint chosen in full view of HeRO&#8217;s split result.</strong> Days-alive-and-ventilator-free is a good composite; it also failed to reach significance, while the mortality signal that did reach significance was fragile. Neither is a template; both are a warning about how much the choice decides in advance.</p></li><li><p><strong>A randomisation design that reckons with contamination.</strong> HeRO randomised individual babies, which is clean statistically but leaky in practice: the same nurses and doctors care for both arms, and a nurse or clinician who learns to read the score cannot unlearn it on the hidden babies, they are much smarter than we think. Randomising whole units, or phasing the score in unit by unit over time, can reduce that cross-arm contamination, but at the cost of far larger samples and vulnerability to centre effects and secular trends. It is a trade-off, not an upgrade.</p></li><li><p><strong>A scale no single centre commands.</strong> HeRO enrolled 3,003 infants across nine units over six years and still produced a mortality estimate whose confidence interval only just excluded no effect. A future trial&#8217;s size cannot simply be read off HeRO; it depends on the endpoint, the baseline risk, the expected effect, the clustering and the adherence, and on whether the intervention is the score alone or the score-plus-response bundle. What is certain is that the arithmetic points away from any one hospital, and our model was built in one hospital, on one hospital&#8217;s data.</p></li><li><p><strong>The onset you still cannot see.</strong> Even a perfectly designed, fully funded trial inherits the unobservable start of sepsis. Measure the operational chain above and the mechanism grows clearer; it does not become transparent.</p></li></ul><p>None of this is a reason not to do it. It is a description of what doing it properly costs, and of why the distance between a good model and a proven pathway is measured in years and thousands of patients rather than in another analysis of the data already in hand.</p><h2>A note on how HeRO reached the bedside anyway</h2><p>Worth tethering back to Episode 5&#8217;s regulatory thread: HeRO did not wait for the trial to reach patients. It entered the trial already 510(k)-cleared, for its intended use as an ECG-derived measure of reduced heart-rate variability and transient decelerations, not for diagnosing sepsis and not for improving survival. That clearance rested, as the 510(k) route does, on substantial equivalence to an existing device and on adequate safety and performance within that stated intended use. </p><blockquote><p><em>A <strong>510(k)</strong> is the main US regulatory route through which many medical devices receive FDA clearance. The name refers to section 510(k) of the Federal Food, Drug, and Cosmetic Act. Rather than proving clinical benefit from scratch, the manufacturer generally has to show that the device is <strong>substantially equivalent</strong> to a legally marketed &#8220;predicate&#8221; device: it has the same intended use and comparable safety and performance. Clearance therefore does not necessarily mean that the device has been shown in a randomized trial to improve patient outcomes.</em></p></blockquote><p>The randomised trial addressed a separate question the clearance never touched: clinical utility, whether showing the number to a clinician at three in the morning changes the night for the better. Regulatory clearance and clinical benefit are different claims, established by different means. Only the benefit claim necessarily calls for a clinical-impact trial, though what regulators demand in practice varies with the device, its intended use and the jurisdiction.</p><h2>The second question, hiding inside the first</h2><p>Even HeRO leaves its own question open. When the smallest survivors were followed up at 18 to 22 months, the composite outcome of death or neurodevelopmental impairment was not significantly reduced, roughly 39% in the displayed group against 44% in controls, a relative risk near 0.87 that did not clear significance (<a href="https://doi.org/10.1016/j.jpeds.2019.12.066">Schelonka et al., 2020</a>). Two caveats matter. The outcome was available for only about 72% of eligible infants, and attrition on that scale limits what can be concluded. And the analysis is of a composite of death or impairment, not of the survivors themselves; asking whether the extra survivors were better off means conditioning on survival, which brings its own selection problem. What can fairly be said is narrower: the follow-up did not show that the mortality benefit was matched by an improvement on the death-or-impairment composite. (A later analysis did find a benefit in the narrower subgroup of extremely preterm infants who developed sepsis, <a href="https://doi.org/10.1016/j.earlhumdev.2021.105419">King et al., 2021</a>, a real but subgroup-level signal a single trial cannot settle.) So the question recurses: the study that suggested displaying the score saves lives raised a further one, did improved survival come with improved, or at least not worse, neurodevelopmental outcomes, that it was not built to answer.</p><h2>Why, fifteen years on, there is still only one</h2><p>Before leaning on the word &#8220;only,&#8221; I went looking for a second, and so had others. The Dutch systematic review already mentioned searched four databases and found fifteen papers on HRC monitoring in preterm infants; across all of them, exactly one randomised trial with clinical outcomes, the HeRO trial, reported three times (<a href="https://doi.org/10.1159/000531118">Koppens et al., 2023</a>). My own search of the trials registry turned up no second completed randomised outcome trial. The nearest newer entry is an observational, model-development pilot rather than a second impact trial (<a href="https://clinicaltrials.gov/study/NCT07254559">NCT07254559</a>).</p><p>One trial and many second looks into the same dataset, the septicaemia-mortality question (<a href="https://doi.org/10.1038/pr.2013.136">Fairchild et al., 2013</a>), the neurodevelopmental follow-up, whether the score performed equally across Black and White infants (<a href="https://doi.org/10.1038/s41390-023-02470-z">Sullivan et al., 2023</a>). It stands alone not because the question stopped mattering, the same review calls plainly for a large international RCT, but because a trial like this is enormous, slow and hard to fund, and because a result positive enough to be quoted let the field feel the question was answered when it was answered only halfway.</p><p>That is the trial that would take. The number can already tell you the baby&#8217;s risk has risen. One large randomised trial suggests that displaying this particular score may reduce mortality, a real and hard-won signal, but a single fragile result, not a settled clinical pathway. What no trial has yet shown, for HeRO and still less for a single-centre model built on a replay, is which move at the bedside, at three in the morning, turns a rising number into a better morning. Until a study is built to answer that, the model is a very good question still waiting for an answer.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Abonneren&quot;,&quot;language&quot;:&quot;nl&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Road to CEPAS 2026! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Typ je e-mailadres&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Abonneren"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>How I used Claude and ChatGPT</h3><p>The argument, the framing and every editorial decision here are mine, including the choice to treat the endpoint correction as a sidebar rather than the spine, to keep the FDA point to a single paragraph, and to fold the neurodevelopmental follow-up in as a second open question rather than let it take over.</p><p>I used Claude to interrogate and cross-check the papers: it surfaced the primary-versus-secondary-endpoint distinction I had been carrying loosely, verified the figures and the systematic review&#8217;s risk-of-bias findings, and ran the registry and literature searches behind the &#8220;only one&#8221; claim, flagging the newer observational entry rather than letting the tidy version stand. The interpretation, and the responsibility for it, are mine.</p><p>ChatGPT made the infographic and reviewed the manuscript for content and readability, removed the em-dashes and proposed technical adaptations.</p><p><em>Next: with the trial mapped, the question is who would ever run it, and what it means to build toward a validated response protocol when the field&#8217;s incentives reward the next model over the missing trial. That is Episode 7.</em></p>]]></content:encoded></item><item><title><![CDATA[The Model Leaves the Lab]]></title><description><![CDATA[Road to CEPAS 2026 &#8212; Episode 5]]></description><link>https://road2cepas.substack.com/p/the-model-leaves-the-lab</link><guid isPermaLink="false">https://road2cepas.substack.com/p/the-model-leaves-the-lab</guid><dc:creator><![CDATA[Quantum Neonatology]]></dc:creator><pubDate>Mon, 29 Jun 2026 18:30:39 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>At the end of the last <a href="https://road2cepas.substack.com/p/the-honest-problem">episode I</a> made a promise. Buried in the discussion section of our own 2023 paper is a reference to European medical-device regulation that I cited and then walked straight past. I said Episode 5 would pick it up and not move on. So here we are.</p><p>Regulatory sentences are easy to walk past because they read like housekeeping. You build a model, you report an AUC, and somewhere near the end a reviewer-pleasing sentence acknowledges that turning this into something used on real babies would mean engaging with European device law. Cite the regulations, nod, move to the conclusion. I have done exactly this. It&#8217;s not dishonest, but it does quietly imply that the regulation is the last, dull step, A lot of paperwork between a finished model and the bedside.</p><p>Yet, it isn&#8217;t the last step. It&#8217;s a different question entirely, and it&#8217;s the one the whole series has been circling.</p><h2>The sentence I should have sat with</h2><p>Our discussion makes an argument I now think is the most important thing in the paper, and it&#8217;s not the AUC. It says, roughly, that to build a medical device you have to specify an intended use <em>together with</em> a performance claim &#8212; and the performance claim has to be good enough to identify any hazard that comes from using the device in practice. </p><p>Read that again with Episode 4 in mind. We spent a whole episode on the alarm policy and the precision problem, the finding that at the operating point we chose, most alarms are not followed by sepsis. In the research frame, that&#8217;s a PPV number, a point on a curve, a thing you note and contextualise. In the regulatory frame it is not a number. It is a <em>hazard</em>. A false alarm that triggers a workup, or antibiotics, or a clinician&#8217;s decision to treat is over-treatment, and over-treatment is precisely the kind of harm the performance claim is supposed to be sensitive enough to catch. The regulation takes the thing we treated as a metric and reclassifies it as a risk to a patient.</p><p>That reframing is the whole episode. Everything below is consequence.</p><h2>Which regulation, and why it matters</h2><p>The first thing to settle is which regulation even applies, because Europe has two and they classify by completely different logic. The in-vitro diagnostic regulation governs devices whose medical purpose depends on examining specimens &#8212; a sample taken from the body, run through a test. The medical device regulation, the MDR, governs everything else, including software that stands on its own.</p><p>The word &#8220;diagnostic&#8221; pulls the reflex toward the in-vitro side, sepsis, blood cultures, diagnosis everywhere. But our model never touches a specimen. It runs on continuously monitored physiology and data already in the record, and produces a score that informs a clinical decision. That is software as a medical device, squarely under the MDR. The in-vitro regulation, despite the diagnostic subject matter, simply isn&#8217;t the instrument that governs a thing like this.</p><p>This isn&#8217;t pedantry. Which regulation you&#8217;re under decides everything downstream, what evidence you owe, whether an independent body has to audit you, how long it takes, what it costs. So: software, under the MDR. What does the MDR do with software that informs a clinical decision?</p><h2>Rule 11, and why this lands high</h2><p>The MDR has a single rule &#8212; Rule 11 &#8212; that classifies software, and it is notorious for pushing almost everything upward. The logic runs by worst credible harm, not by how often the software is right. Software that provides information used in diagnostic or therapeutic decisions starts at Class IIa. If the decisions it informs could cause death or irreversible harm, it&#8217;s Class III. If serious deterioration, Class IIb. And there&#8217;s a second limb: software that monitors vital physiological parameters, where the <em>nature of the variation</em> in those parameters could put the patient in immediate danger, lands at IIb.</p><p>Walk our model through that. It monitors continuous vital-sign physiology. The variation it watches for is the early signature of a baby becoming septic, a state where deterioration can be immediate and severe. The decision it informs is whether to start antibiotics in a preterm infant, where error in either direction has consequences. On a conservative reading, this is not low-risk software. Class IIb is a very plausible landing point; depending on the exact intended use and performance claim, a Class III argument is not impossible.</p><p>Class IIb is not Class I. Class I you can largely self-declare. From IIa upward an independent notified body audits your quality system and reviews your technical file before anything is allowed near a patient. The thing that pushes you over that line is not the sophistication of the model. It&#8217;s the severity of the decision it touches. A simpler, more interpretable model with the same intended use lands in exactly the same class, a point that should sound familiar to anyone who read Episode 3.</p><h2>&#8220;It works in a simulation&#8221; is the start of the conversation</h2><p>There&#8217;s a story that should be taught alongside every sepsis-prediction paper, and it isn&#8217;t ours. The <a href="https://jamanetwork-com.utrechtuniversity.idm.oclc.org/journals/jamainternalmedicine/fullarticle/2781307">Epic Sepsis Model</a> was built into one of the most widely used electronic record systems in the world and deployed across hundreds of American hospitals, about as far from &#8220;research artifact&#8221; as a model can get. Then an external team at Michigan validated it independently and found it missed roughly two-thirds of sepsis cases at the threshold in use. Deployed everywhere; externally validated almost nowhere until someone finally checked.</p><p>That gap &#8212; between <em>deployed</em> and <em>validated</em> &#8212; is the gap the regulation exists to close. And it&#8217;s the gap our own staged plan, the one I cited approvingly in the paper, is built around. We named the path: prospective validation first, then a clinical validation study, then the hard outcomes, mortality, sepsis episodes, antibiotic days. That ladder isn&#8217;t bureaucratic decoration. Each rung is a different claim about the world, and the regulation is essentially asking you to climb it in order and show your work at each step. A retrospective AUC, however good, is a claim about a dataset. It is not yet a claim about a baby.</p><p>The model in this lineage that actually climbed the ladder is the old one. The heart-rate-characteristics work I traced back in Episode 1 ran a genuine <a href="https://www.jpeds.com/article/S0022-3476%2811%2900671-8/abstract">prospective randomised trial</a> &#8212; 3003 very-low-birth-weight infants across nine NICUs, with a reduction in mortality when the monitor was displayed, published in 2011. Fifteen years later it remains the canonical prospective randomised trial in this literature. Much of the field since, ours included, has stopped at the rung marked &#8220;good ROC curve.&#8221; The regulation is, in a sense, just the formal version of the question that trial answered and the rest of us have been deferring: <em>does using this change what happens to the patient?</em></p><h2>The ground is moving while I write this</h2><p>I should be honest that the regulatory picture in 2026 is not settled, and a research diary should say so rather than pretend the law is a fixed monument. The AI Act now adds a second layer: an AI medical device that requires notified-body assessment will generally fall into the high-risk AI category, which our model, on the reading above, would. That brings requirements around data governance, representativeness of the training data, transparency, human oversight, logging, and post-market monitoring. The direction of travel is toward avoiding two wholly separate assessments and integrating the AI Act requirements into the existing product-conformity architecture where possible, rather than running them in parallel. The exact timing and mechanics remain in flux: the high-risk AI timelines and a December 2025 proposal to simplify the MDR/IVDR framework, which secondary commentary reads as touching the software-classification logic, including Rule 11, are still moving through the system.</p><p>So the exact procedural path may move, and the specific class boundaries with it. What won&#8217;t shift is the underlying logic: severity of the decision drives the obligation, and deployment is gated on evidence the field has mostly not produced.</p><p>I&#8217;m flagging this partly for accuracy and partly because it matters for the talk in Lyon. The details will have moved again by October. The argument won&#8217;t.</p><h2>Back to 3am</h2><p>I ended the last episode at a bedside: a score crosses a threshold in front of a clinician at three in the morning, and the only question that matters is whether that becomes earlier treatment or just one more thing beeping.</p><p>The regulation is what stands between those two outcomes, and I&#8217;ve spent this episode learning that I&#8217;d been treating it as the wrong kind of thing. Not paperwork after the science. A second, harder science &#8212; the one that asks not &#8220;can the model tell true from false&#8221; but &#8220;when this fires at 3am and someone acts on it, is the baby better off.&#8221; Our paper has an answer to the first question and, honestly, an IOU on the second.</p><p>Next episode I want to follow that IOU somewhere specific: what a real clinical validation of a model like ours would actually have to look like, and why almost nobody has done one.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Abonneren&quot;,&quot;language&quot;:&quot;nl&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Road to CEPAS 2026! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Typ je e-mailadres&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Abonneren"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>How I used Claude</h3><p>The regulatory reading in this episode is not my home turf, and that shaped the division of labour. I asked Claude to pull the current state of the EU device rules &#8212; the MDR classification logic, the software rule, the in-vitro versus medical-device distinction &#8212; and the 2026 developments around the AI Act, because this is exactly the kind of material that dates fast and that I&#8217;d otherwise half-remember from a conference slide. Claude surfaced the Rule 11 mechanics, the Epic Sepsis Model external-validation figure, and the live status of the AI Act overlay and the late-2025 simplification proposal.</p><p>The argument is mine. The decision to open on our own paper&#8217;s &#8220;intended use plus performance claim&#8221; sentence, and to read the false-alarm rate from Episode 4 as a regulatory <em>hazard</em> rather than a metric, was the spine I wanted before any research happened; the regulation just gave it teeth. I had Claude check the Rule 11 classification reasoning against current guidance rather than take my reconstruction of it on trust, because a confidently wrong regulatory class is worse than no class. Where the law is genuinely unsettled in 2026, I&#8217;ve tried to say so plainly rather than launder uncertainty into authority &#8212; and that paragraph got read adversarially before it went in.</p><p><em>Next: Episode 6 &#8212; what a real clinical validation of a model like this would have to look like, and why, fifteen years on, that 2011 trial still stands largely alone.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="6000" height="4000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4000,&quot;width&quot;:6000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;white and black hallway&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="white and black hallway" title="white and black hallway" srcset="https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1566010773563-954336289897?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw0MXx8bGFifGVufDB8fHx8MTc4MjMyOTkzMnww&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@hikeshaw">Bofu Shaw</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div>]]></content:encoded></item><item><title><![CDATA[Episode 4 — Our Own Work, Warts and All]]></title><description><![CDATA[Road to CEPAS 2026 &#183; Episode 4]]></description><link>https://road2cepas.substack.com/p/episode-4-our-own-work-warts-and</link><guid isPermaLink="false">https://road2cepas.substack.com/p/episode-4-our-own-work-warts-and</guid><dc:creator><![CDATA[Quantum Neonatology]]></dc:creator><pubDate>Wed, 17 Jun 2026 17:01:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!GNNg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GNNg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GNNg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!GNNg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!GNNg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!GNNg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GNNg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74598797-ca11-45f1-b045-383012f478bd_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1785585,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://road2cepas.substack.com/i/200793180?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GNNg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!GNNg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!GNNg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!GNNg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74598797-ca11-45f1-b045-383012f478bd_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>For three episodes I have been reading through other people&#8217;s papers. I mapped the field, then read four of its anchor models side by side and argued that the two decisions that matter most in any of these systems, which signals you give the model, and how you turn its scores into alarms, are the two decisions almost no one documents clearly.</p><p>It would be too easy to leave the argument there. The most credible place to test it is a paper I cannot hide behind a citation: my own.</p><p>So this episode turns the same reading onto <a href="https://pubmed.ncbi.nlm.nih.gov/37369173/">van den Berg et al. 202</a>3, the WKZ Utrecht late-onset sepsis model I am senior author on. The same two decisions &#8212; signal stack and alarm policy &#8212; with the candour raised, because I owe it to you that after the last episode, and because the limitations of our own paper point directly to where this series goes next.</p><p>I will state the obvious first, so it is not mistaken later for false modesty: I think it is good work. It was done by incredibly smart people from our <a href="https://3ai.umcutrecht.nl/pillars/implementation/ai-for-health/">AI for health</a> department and many more before that in the ADAM program. It is the largest unselected single-centre cohort in this corner of the literature, and it carries the most extensive retrospective impact simulation anyone has done for LOS prediction to date. I stand behind both claims. But &#8220;good work with real limitations&#8221; has two halves, and the second is the reason for this episode.</p><h2>The number we didn&#8217;t put in the abstract</h2><p>The abstract reports that the algorithm detects LOS in at least 47% of patients before clinical suspicion, without exceeding an alarm-fatigue threshold of three alarms per day. That is the sentence most readers take away, and the one most likely to be quoted.</p><p>At that same operating point, the precision is 3.93% on the training set and 4.41% on the test set. This number is not in the abstract.</p><p>A precision of 4% means that of every hundred alarms the system raises, roughly ninety-six are false, and not at some reckless sensitivity setting, but at the operating point we chose precisely because it matched the alarm burden a NICU already tolerates. The defence the paper offers is genuine and I still hold to it: three alarms a day is approximately how often the WKZ team already considers sepsis during routine care, so the system adds little net burden even when most alarms do not pan out. We did a small (unpublished) study where we found that 80% of the documented sepsis evaluations did not result in a blood culture. A bedside clinician already works with a high base rate of &#8220;check, probably nothing.&#8221; </p><p>Still, the asymmetry in the writing is worth naming. A detection fraction went into the abstract; the false-alarm rate stayed in Table 3. They come from the same alarm configuration, but they summarize different aspects of it: the fraction of LOS episodes flagged before the culture, and the positive predictive value of the alarms generated by that policy. We foregrounded the more flattering summary. That is not data manipulation &#8212; every figure is in the paper for anyone who reads as far as the table &#8212; but it is a framing decision, and an honest account of the paper has to treat it as one. If &#8220;prediction is not utility&#8221; means anything, it means the 4% deserves at least as much prominence as the 47%.</p><h2>Signal stack: the strength and the limitation are one decision</h2><p>The first of the two decisions is what the model is allowed to see.</p><p>Heart rate and oxygen saturation, and nothing else. Low-frequency, one sample a minute, the values that scroll across every monitor on the ward. We deliberately set aside the inputs that would have raised the AUC: CRP, blood pressure, white cell counts. The reason was principled. A clinician orders a CRP only when they already suspect something, so a model trained on CRP learns to detect clinical suspicion and then claims credit for predicting it. We were strict about excluding anything carrying that prior knowledge, and I would defend that strictness without hesitation. It is the best decision in the paper.</p><p>It is also the source of two of the paper&#8217;s hardest limitations, and the point worth dwelling on is that the strength and the limitations are not separate items to be weighed against each other. They are one decision seen from two sides.</p><p>Because the model keys on physiological deterioration alone, it cannot say why a baby is deteriorating. The paper states this directly: the algorithm may not be specific to sepsis and could also respond to other inflammatory conditions such as NEC. A model built on the body&#8217;s distress signal will respond to distress, whatever its cause. We call it a sepsis model; it is more accurately a deterioration model aimed at a sepsis question.</p><p>The same minimalism is part of why the precision is so low. With two physiological streams and no specific marker, there is a hard ceiling on how cleanly the model can separate the baby who is becoming septic from the one who is briefly unstable for any of a dozen ordinary reasons. The interpretable seven-feature logistic regression is not underpowered for want of a better architecture, the supplement shows XGBoost and GAMs reach the same performance, with no significant difference between them. It sits at that ceiling because the signal stack we chose, for good reasons, cannot see past it. The principled choice and the (disappointing) precision are the same fact stated twice.</p><h2>Alarm policy: an operational convenience that became a convention</h2><p>The second decision is the one I now think we explained least, and it is the one Episode 3 identified as under-documented across the whole field.</p><p>A model produces a score every hour. A score is not an alarm. Turning one into the other requires a policy, and ours had three parts:</p><ul><li><p><strong>A refractory period.</strong> Once an alarm fires, further alarms are silenced for a set window, so a stably high-risk baby does not alarm every hour. We set that window to eight hours, on the stated reasoning that eight hours is the length of a nursing shift at the WKZ NICU.</p></li><li><p><strong>A true-positive window.</strong> An alarm counts as catching the sepsis if it falls anywhere from 24 hours before to 12 hours after the blood culture.</p></li><li><p><strong>A multi-threshold escalation rule</strong>, in which a higher-risk alarm can break through the silence imposed by a lower-risk one.</p></li></ul><p>The justification for the refractory period is the revealing part. Eight hours, because that is a shift. This is an operational convenience, a reasonable one for a working ward, but it is not a clinically derived parameter. No one demonstrated that eight hours is where early detection and alarm fatigue trade off best. It is where the staff rotation happens to fall.</p><p>Whether that matters is something the supplement lets us check, and the answer is that it does, moderately. Varying the refractory period to 4 or 12 hours (Tables S7 and S8) moves recall and precision, not dramatically, but enough to show that the headline figure depends on a number chosen for rostering reasons. The true-positive window matters more. Our default window spans 36 hours and counts an alarm up to 12 hours after the culture as a successful early warning. Restrict it to the stricter interval from 12 hours before the culture to the culture itself and multi-alarm recall falls from the high-50s and 60s into the high-30s to mid-50s (Table S3). The &#8220;at least 47%&#8221; is real, but it rests on a generous definition of what counts as catching the episode in time.</p><p>This is more than an internal quibble because the policy did not stay internal. Episode 3 documented that this alarm framework, the eight-hour silencing, the multi-threshold escalation, the alarms-per-patient-day accounting, was taken up by other groups. Yang 2024 cites our paper and reuses both the eight-hour period and the escalation logic; Meeus echoes the alarm-burden accounting. I presented that propagation in Episode 3 as a point of pride. Seen from the other direction, though, a refractory period justified by our shift pattern has propagated into other people&#8217;s models as a convention. The element we documented least rigorously is the one the field borrowed most directly. That ought to make all of us  cautious about how much of a shared methodology is in fact a shared assumption.</p><h2>The proxy at the centre of everything</h2><p>Beneath both decisions lies a single substitution that the whole impact story depends on.</p><p>We do not know the moment a clinician first suspected sepsis; it is not reliably recorded. So we used the timestamp of the blood culture as a stand-in for that moment &#8220;t = 0&#8221;. Every claim about detecting LOS before clinical suspicion is, more precisely, a claim about detecting it before the blood culture was drawn.</p><p>The paper concedes this directly: because actual suspicion could arise earlier, using the culture time as the proxy can lead to an overestimation of the early-warning performance. It is the cleanest limitation in the paper. The entire early-warning proposition rests on a proxy the authors themselves flag as probably flattering. If a clinician was already uneasy an hour before drawing the culture, part of our measured lead time is an artifact of when the culture was recorded, not of when the model spoke first.</p><p>There is a smaller version of the same problem in the data. A number of blood cultures arrived timestamped at exactly midnight, an artifact of records where the time was missing and defaulted to 00:00. We imputed those to noon, closer to morning rounds, less likely to be far off. It is a defensible fix, and we confirmed that dropping the contaminated records did not change the results. But it means that a portion of our t = 0 anchors, the reference point the whole simulation turns on, are reasoned estimates of what the clock read.</p><p>While re-reading the supplement for this episode I also found something I cannot fully explain. The train/test baseline table reports a CRP-above-10mg/L rate in the control group of roughly 80%, against 8.5% in the main-text table &#8212; a figure that cannot be right for a control group, and one the page appears to print twice with differing control numbers. I cannot reconstruct from the published PDF exactly what went wrong in producing that table. It is probably proofing fault, but I would be a poor critic of my own paper if I only noticed the errors that flatter it.</p><p>One further definitional point quietly inflates the false-alarm count. We labelled only culture-positive cases as sepsis. Culture-negative &#8220;clinical sepsis&#8221;, babies treated as septic who never grew an organism, sits in the control group. When the model fires on one of them it is scored as a false positive, even where a clinician might have agreed with it. So some unknown fraction of the 96% of &#8220;false&#8221; alarms may not be false. That cuts in our favour, but a muddied outcome definition is muddied whichever way it leans.</p><h2>Simulation is not impact</h2><p>This brings me to the claim I am at once most proud of and most wary of. We described our longitudinal alarm analysis as the most extensive clinical impact assessment in the LOS prediction literature to date.</p><p>It is worth being clear about what kind of object it is. It is a retrospective simulation run on historical data, the model applied to records of babies whose outcomes were already settled, asking what the alarms would have done. No clinician saw a score. No decision changed. Nothing that happened to those babies happened because of this model.</p><p>The paper&#8217;s final move is its most honest. Listing what a real evaluation would require, the authors note that counterfactuals &#8212; would this sepsis have been caught later without the model? &#8212; are not measured. That is the thesis of this entire series, stated in our own paper, in our own words, two years before I began writing these episodes. We demonstrated prediction. We simulated impact. We did not show utility, and we said so.</p><p>That gap, between the most extensive impact simulation and actual clinical impact, is not a flaw in the paper. It is the frontier at which the paper honestly stops. Every model in Episode 2&#8217;s field map stops at roughly the same line. The open question is no longer whether we can predict LOS early; the field has answered that many times over, ours included, within an AUC band everyone now clusters in. The question is what happens when one of these systems is placed in front of a clinician at three in the morning and a score crosses a threshold, whether the early warning becomes earlier treatment, or simply one more thing beeping.</p><p>That is deployment. And deployment, in Europe, is where the model stops being a research artifact and becomes something the regulator has views on. Berg 2023 gestures at this already: two footnotes in the discussion point to the EU Medical Device and In-Vitro Diagnostic Regulations, almost in passing, as the frame to which any performance claim would eventually have to answer.</p><p>We dropped those two footnotes and moved on. Episode 5 will pick them up and not move on.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Abonneren&quot;,&quot;language&quot;:&quot;nl&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Road to CEPAS 2026! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Typ je e-mailadres&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Abonneren"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>How I used Claude</h2><p>For this episode the work was reading rather than searching. I gave Claude the Berg 2023 paper and its supplementary materials and asked it to build the critique against the primary source instead of from my memory of my own paper. That was the right instinct, because reading the supplement closely brought back things I had half-forgotten. The sensitivity of the headline recall to the true-positive window, the comparison in Table S3 between the strict and generous intervals, is set out in the supplement, and having it in front of me again sharpened the alarm-policy section. Claude also pointed out the implausible control CRP figure in the supplementary baseline table, which I had never noticed.</p><p>The structural decision, to organise around the two decisions from Episode 3 rather than re-run all eight dimensions, was mine, and I want to be exact about that, because this section is the easiest place in the series to quietly overstate what the tool did. Claude set out both options and argued for the two-decisions structure; I chose it because we had already done the eight-dimension grid in Episode 3 and repeating it would have read as padding. The judgements about what counts as a real limitation versus a defensible choice are mine, including the uncomfortable one about what we put in the abstract.</p><p><em>Next: Episode 5 &#8212; the model leaves the lab. What European medical-device regulation actually requires of a system like this, and why &#8220;it works in a simulation&#8221; is where the conversation begins rather than ends.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/p/episode-4-our-own-work-warts-and/comments&quot;,&quot;text&quot;:&quot;Laat een reactie achter&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://road2cepas.substack.com/p/episode-4-our-own-work-warts-and/comments"><span>Laat een reactie achter</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Episode 3: A Field Full of ROC Curves]]></title><description><![CDATA[Similar, but not the same]]></description><link>https://road2cepas.substack.com/p/episode-3-a-field-full-of-roc-curves</link><guid isPermaLink="false">https://road2cepas.substack.com/p/episode-3-a-field-full-of-roc-curves</guid><dc:creator><![CDATA[Quantum Neonatology]]></dc:creator><pubDate>Wed, 03 Jun 2026 18:01:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ntn_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ntn_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ntn_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic 424w, https://substackcdn.com/image/fetch/$s_!Ntn_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic 848w, https://substackcdn.com/image/fetch/$s_!Ntn_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic 1272w, https://substackcdn.com/image/fetch/$s_!Ntn_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ntn_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:294624,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/heic&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://road2cepas.substack.com/i/199325062?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ntn_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic 424w, https://substackcdn.com/image/fetch/$s_!Ntn_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic 848w, https://substackcdn.com/image/fetch/$s_!Ntn_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic 1272w, https://substackcdn.com/image/fetch/$s_!Ntn_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7ce9772-7465-4f06-8ae0-528cc6ec3d5f_1672x941.heic 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In <a href="https://road2cepas.substack.com/p/episode-2-mapping-the-field">Episode 2</a>, I mapped the field and came away with three claims. One of them was technical convergence, across every group doing continuous-physiology machine learning for late-onset sepsis in preterm infants, the headline AUCs land between 0.78 and 0.88. I called the convergence real and said it wasn&#8217;t a methodological artefact.</p><p>I want to take that back a little. Not because the numbers are wrong, they&#8217;re right. But because lining four AUCs up in a row and calling them a convergence quietly implies they are four readings of the same thing. They are not. This episode is the close read that shows why and it lands on a more useful question than &#8220;how high is the AUC.&#8221;</p><p>I promised at the end of Episode 2 that this one would go deep on two decisions: <strong>the signal stack</strong> and <strong>the alarm policy</strong>. I&#8217;m keeping that promise, because after reading the four anchor papers side by side, I&#8217;m convinced those two decisions matter more than anything in the modelling and they are exactly the two things we need to discuss.</p><p>The four papers, for reference: our own from WKZ Utrecht (<a href="https://doi.org/10.1016/j.compbiomed.2023.107156">Berg 2023</a>), the multi-NICU study from UVA (<a href="https://doi.org/10.1038/s41390-022-02444-7">Kausch 2023</a>), the Eindhoven/M&#225;xima paper (<a href="https://doi.org/10.1016/j.cmpb.2024.108335">Yang 2024</a>), and the Antwerp/Innocens paper (<a href="https://doi.org/10.1016/j.jpeds.2023.113869">Meeus 2024</a>). Full eight-dimension comparison is on the <a href="https://neonatology.ai/sepsis/episode-03">companion page</a>; here I&#8217;ll walk the argument.</p><h2>The first decision: which signals, how fast, from what</h2><p>Before any machine learning happens, you choose a signal stack. It sounds like a technicality. It is the single most consequential design decision in the whole field, and it&#8217;s really three decisions wearing one coat: which signals, sampled how fast, from which device.</p><p>Two of the four papers contain a finding on this that I think is a bit underrated, and in both cases it&#8217;s a sentence or two that the paper itself doesn&#8217;t dwell on.</p><p>The first is from Kausch. They built their model, POWS, and then asked whether it mattered if the heart rate came from an ECG or from a pulse oximeter. It didn&#8217;t. Performance was within about 0.01 AUC either way. That sounds like a footnote. It isn&#8217;t. It means a continuous-physiology sepsis model can run on a standalone pulse oximeter, no ECG leads, no specialised monitor. For a well-resourced NICU that&#8217;s a convenience. For everywhere else, it&#8217;s the difference between a deployable tool and a research curiosity. The most important sentence in the paper for global deployment is a sensitivity analysis.</p><p>The second is from Yang, they ran a sampling-rate ablation: the same pipeline fed raw waveforms, then 1 Hz vitals, then one sample per minute, then one per hour. Raw waveforms scored 0.886. One Hertz scored 0.875, basically the same. One sample per minute, 0.825. One per hour, 0.687, the signal essentially gone. Two lessons fall out of that. You do not need waveform processing; minute-resolution numerical data, the kind any modern monitor already exports, gets you almost all the way. But once you summarise to hourly aggregates, the signal collapses. So there is probably a floor, minute-by-minute, and it is now empirically settled where it sits.</p><p>Put those together and the signal-stack picture is clearer than the AUCs suggest. More channels is not better. Yang uses three signals, Kausch two, and they bracket the field. Meeus uses seven, heart rate, respiration, oxygen saturation, perfusion index, temperature, FiO&#8322;, glucose, and it does not buy a cleanly higher comparable number. What moves performance is sampling resolution and a well-chosen feature like Kausch&#8217;s heart-rate&#8211;oxygen cross-correlation. Not the length of the signal list.</p><h2>The AUCs are not measuring the same thing</h2><p>Here are the four AUC headline numbers. Berg: 0.79. Kausch: about 0.79 on external validation. Yang: 0.875. Meeus: 0.93, statistically significant, for what that matters. In Episode 2 I let those sit together as a band. Read the papers and the band falls apart, not because any number is wrong, but because each one answers a different question.</p><ul><li><p><strong>Berg&#8217;s 0.79</strong> is a cross-sectional AUC at the moment of clinical suspicion, on a whole-NICU population of every infant 32 weeks or under, with culture-positive sepsis as the outcome. (For honesty&#8217;s sake: our cross-sectional figure is 0.73 on the training set and 0.79 on test &#8212; at or below the floor of the band I quoted in Episode 2.)</p></li><li><p><strong>Kausch&#8217;s 0.79</strong> is a prediction of sepsis within the next 24 hours, on a VLBW-only cohort, a narrower, sicker slice, and crucially it is an <em>external</em> number, the model tested at hospitals it was not trained on. (!)</p></li><li><p><strong>Yang&#8217;s 0.875</strong> is measured 6 hours before the sepsis workup, on a cohort of 119 infants that was deliberately cleaned to clear cases and clear controls.</p></li><li><p><strong>Meeus&#8217;s 0.93</strong> is for a model predicting late-onset sepsis <em>and</em> necrotising enterocolitis jointly, not LOS alone, and it is computed on cross-validated time windows, a metric the Meeus paper itself flags as optimistic on heavily imbalanced data.</p></li></ul><p>Four different outcomes, or horizons, or cohorts, or evaluation units. The numbers being close is a coincidence of four separate measurements, not agreement between four readings of one quantity. &#8220;Similar, but not the same&#8221; is the honest description, and it&#8217;s why the subtitle of this episode is what it is.</p><p>So if AUC won&#8217;t carry a fair comparison, what will? The metric that survives translation between these papers is a clinical one: <strong>what fraction of sepsis episodes does the model catch </strong><em><strong>before</strong></em><strong> the clinician does</strong>. On that metric the picture is consistent. Berg detects roughly 47 to 60 percent of patients before clinical suspicion. Meeus catches 69 percent of all episodes and 81 percent of severe ones, with a median head start around 10 hours. Kausch sees risk climb significantly 23 to 24 hours ahead. Yang reports 96 percent, but patient-wise, on 119 infants, with a specificity of 19 percent, so it belongs in a different bucket than the others.</p><p>Set Yang&#8217;s small-cohort outlier aside and the field converges on something like <strong>half to two-thirds of late-onset sepsis episodes detectable hours before clinical recognition.</strong> That is the real state of the art. It is a less impressive number than &#8220;AUC 0.9,&#8221; and a far more honest one. It is also the number a clinician can actually reason about.</p><h2>The second decision: the alarm policy</h2><p>Say you have a model. It emits a risk score every hour for every baby on the unit. When does a clinician get told? That rule, the alarm policy, is what converts an AUC into something a human being at a bedside actually experiences. And it is chosen, across these four papers, in four genuinely different ways, mostly without much justification for why.</p><ul><li><p><strong>Berg</strong> generates hourly scores and sets thresholds to fixed unit-wide alarm rates. A new alarm is muted for an 8-hour refractory period, the length of a nursing shift at our NICU, unless a <em>higher</em> threshold is crossed, in which case it escalates. False alarms are reported as alarms per patient-day.</p></li><li><p><strong>Kausch</strong> uses a lighter rule: an alert switches on at a threshold and stays on until 24 hours pass with no further crossing. No refractory period, no escalation tiers.</p></li><li><p><strong>Yang</strong> adopts Berg&#8217;s framework more or less wholesale, same 8-hour silencing, same multi-threshold escalation.</p></li><li><p><strong>Meeus</strong> aggregates hourly predictions into 24-hour buckets and reports alarm-days per week, which echoes Berg&#8217;s alarm-burden accounting without the refractory machinery.</p></li></ul><p>Here is the point that matters. The detection-fraction headlines I just quoted, Berg&#8217;s 47 percent, Yang&#8217;s 96 percent, are not, mostly, a gap between models. They are a gap between alarm policies and cohorts. Change the refractory rule, the threshold scheme, the true-positive window, and the same underlying model produces a very different headline. The alarm policy is the highest-leverage decision in turning a model into a deployable tool, and it is the one most often relegated to a methods paragraph and never defended.</p><p>I&#8217;ll claim one thing for our group here, and leave the rest for a later episode. The alarm-fatigue framework from the Berg paper, hourly prediction, multi-threshold escalation, an 8-hour shift-length refractory period, false-alarm burden counted as alarms per patient-day, has quietly become a small convention. Yang adopted it directly. Meeus echoed its accounting. The modelling across the four groups is roughly equivalent; what propagated was a way of <em>evaluating</em> a model for deployment. That&#8217;s worth saying plainly. Whether our own paper holds up to a harder look is a question for Episode 4.</p><h2>What this leaves us with</h2><p>The two decisions that most determine whether one of these models could ever run at a bedside, what signals you feed it, and how you decide when to alarm, are precisely the two decisions the literature documents least clearly. The modelling, by contrast, is well described, carefully compared, and largely settled: across these papers, architecture barely matters once the features are right.</p><p>That inversion is this episode&#8217;s reason to exist. We have spent fifteen years getting very good at the part that turned out not to be the bottleneck, and we write almost casually about the two parts that are.</p><p>There is a coincidence at the centre of this field that looks like consensus. Four ROC curves, similar but not the same. The convergence is real as an observation and misleading as a measurement, and the more useful question was never &#8220;how high is the AUC&#8221; but &#8220;caught before the clinician, how often, on whom, under what alarm rule.&#8221; From here the series turns toward deployment, and toward the European regulatory frame that decides whether any of this reaches a patient.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Abonneren&quot;,&quot;language&quot;:&quot;nl&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Road to CEPAS 2026! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Typ je e-mailadres&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Abonneren"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>How I used Claude</h3><p>This episode is a four-paper close read, and Claude did the structural work of it with me. I uploaded the information form the four anchor papers, plus from the three supporting-cast papers from the Episode 2 map, and we built an eight-dimension comparison across them: cohort design, signal stack, ML methodology, validation rigour, performance, alarm policy, honest limitations, unique contributions. That comparison grid is what this post is distilled from, and it&#8217;s published in full on the <a href="https://neonatology.ai/sepsis/episode-03">companion page</a>.</p><p>Two moments are worth reporting honestly. First, Claude caught that the handover note I started from described work, a finished comparison sitting in an earlier chat, that it then couldn&#8217;t actually find, it wasn&#8217;t published yet, it stopped and said so rather than inventing it. We rebuilt the analysis from the papers instead. </p><p>Second, I asked it to check the licensing of the four PDFs before we used them, which are open access, which are all-rights-reserved, and it worked through them paper by paper. The short version that came out of that: discussing and comparing papers is unconstrained, redistributing them is not. That&#8217;s why this series links to DOIs and hosts no PDFs.</p><p>It also caught a detail I&#8217;d have glossed: our own Berg paper&#8217;s AUC is 0.73 on training and 0.79 on test, at or below the floor of the &#8220;0.78&#8211;0.88 band&#8221; I quoted last episode. That&#8217;s not an embarrassment to bury. It&#8217;s the cleanest single piece of evidence for this episode&#8217;s whole argument.</p><div><hr></div><p><em>This post is part of &#8220;Road to CEPAS 2026,&#8221; a transparent research diary as I prepare a 20-minute talk for the CEPAS 2026 congress in Lyon. <a href="https://road2cepas.substack.com/p/the-honest-problem">Episode 1</a> traced the HeRO/Moorman lineage; <a href="https://road2cepas.substack.com/p/episode-2-mapping-the-field">Episode 2</a> mapped the field. Episode 4 turns the same close-reading on our own work &#8212; Berg 2023, warts and all. The full eight-dimension comparison and methodology log for this episode is at <a href="https://neonatology.ai/sepsis/episode-03">neonatology.ai/sepsis/episode-03</a>.</em></p><p><em>If I&#8217;ve misread one of these four papers, or weighted a comparison wrongly, tell me &#8212; the whole point of doing this transparently is to fix what I get wrong.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/p/episode-3-a-field-full-of-roc-curves/comments&quot;,&quot;text&quot;:&quot;Laat een reactie achter&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://road2cepas.substack.com/p/episode-3-a-field-full-of-roc-curves/comments"><span>Laat een reactie achter</span></a></p>]]></content:encoded></item><item><title><![CDATA[Episode 2: Mapping the Field]]></title><description><![CDATA[What a transparent literature search with Claude revealed about late-onset sepsis prediction in preterm infants &#8212; and what it means for CEPAS 2026.]]></description><link>https://road2cepas.substack.com/p/episode-2-mapping-the-field</link><guid isPermaLink="false">https://road2cepas.substack.com/p/episode-2-mapping-the-field</guid><dc:creator><![CDATA[Quantum Neonatology]]></dc:creator><pubDate>Fri, 22 May 2026 17:00:57 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3992" height="2242" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2242,&quot;width&quot;:3992,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;aerial view of asphalt road surrounded by trees&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="aerial view of asphalt road surrounded by trees" title="aerial view of asphalt road surrounded by trees" srcset="https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1498637841888-108c6b723fcb?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyMXx8cm9hZHxlbnwwfHx8fDE3Nzg4ODA3OTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@jerrykavan">Jerry Kavan</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>I&#8217;m senior author on one of the more recent ML papers for late-onset sepsis prediction in preterm infants. I collaborate with several of the others doing this work. I thought I knew the field, kept up with literature.</p><p>In <a href="https://road2cepas.substack.com/p/the-honest-problem">Episode 1</a> I traced the lineage from HeRO and the Moorman trial through to the present. The question I&#8217;m building toward for CEPAS 2026 is harder: <strong>where does the field actually stand right now?</strong> Not where the headlines say it is. Where the literature actually says it is.</p><p>What follows is the result of a staged literature search with Claude. Five search rounds, roughly 200 abstracts triaged, full text on seven anchor papers, two systematic reviews. The chat log is the source material; this post is what came out the other end.</p><h2>First surprise: the field has its own dialect</h2><p>I started where any clinician would. PubMed, <em>&#8220;neonatal sepsis&#8221; AND &#8220;machine learning&#8221;</em>, filtered for the last 15 years. 101 papers. The top of the list looked sensible, Kainth&#8217;s 2024 systematic review, a handful of clinical-features models, the usual review pieces.</p><p>Then I narrowed to <em>&#8220;late-onset sepsis&#8221; AND &#8220;preterm&#8221; AND ML</em>. Forty papers. Four obvious anchors: Berg, Kausch, Yang, Meeus. I thought I had the field.</p><p>Three more searches to be sure and the actual count was at least seven academic groups, plus a second systematic review I hadn&#8217;t seen.</p><p>The problem was in the vocabulary. My searches were in <strong>machine learning terminology</strong> &#8212; &#8220;machine learning&#8221;, &#8220;prediction model&#8221;, &#8220;algorithm&#8221;. But a substantial part of the field uses <strong>physiology terminology</strong> &#8212; &#8220;heart rate characteristics&#8221;, &#8220;HRC index&#8221;, &#8220;HRV&#8221;, &#8220;visibility graph&#8221;. Those papers don&#8217;t reliably surface on ML keyword searches, even though they&#8217;re doing the same job for the same patients.</p><p>The clearest evidence was the discovery of <strong>two separate systematic reviews of the same clinical problem</strong>, written within months of each other, with almost no overlap in cited literature:</p><ul><li><p><strong><a href="https://doi.org/10.1097/INF.0000000000004409">Kainth et al. 2024</a></strong> (<em>Pediatr Infect Dis J</em>): 19 studies, pooled AUC 0.94. Reviews ML models using clinical and laboratory features. <em>Explicitly excludes</em> vital-signs-only studies.</p></li><li><p><strong><a href="https://doi.org/10.1159/000531118">Koppens et al. 2023</a></strong> (<em>Neonatology</em>, Amsterdam UMC): 15 studies, 8,230 infants. Reviews HRC monitoring for LOS in preterm. <em>This is the vital-signs branch.</em></p></li></ul><p>The two reviews do not cite each other. They are reviewing the same clinical problem, early detection of LOS in preterm infants, through two different methodological lenses, and the camps don&#8217;t talk. Calling them <strong>Branch A (continuous vital signs)</strong> and <strong>Branch B (snapshot clinical and lab features)</strong> is the best way to keep them straight. Our paper is Branch A.</p><p>Searching the right field requires speaking its dialect. That was the first lesson, and one I wish I&#8217;d known when I started <a href="https://road2cepas.substack.com/p/the-honest-problem">Episode 1</a>.</p><h2>The complete map</h2><p>After five staged searches and a careful read of both systematic reviews, the active groups in Branch A, continuous-physiology ML for LOS in preterm infants, are:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bfgf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bfgf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png 424w, https://substackcdn.com/image/fetch/$s_!bfgf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png 848w, https://substackcdn.com/image/fetch/$s_!bfgf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!bfgf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bfgf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png" width="1404" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1404,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:174280,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://road2cepas.substack.com/i/198136663?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bfgf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png 424w, https://substackcdn.com/image/fetch/$s_!bfgf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png 848w, https://substackcdn.com/image/fetch/$s_!bfgf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!bfgf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17aa2dde-0307-497a-b9de-31f5f76541dd_1404x1048.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><ul><li><p><strong>UVA / Med Predictive Science</strong> (US) &#8212; Kausch SL et al. 2023, <em>Pediatric Research</em>. <a href="https://doi.org/10.1038/s41390-022-02444-7">DOI: 10.1038/s41390-022-02444-7</a>. The HeRO heritage; only multi-NICU external validation in the field.</p></li><li><p><strong>WKZ Utrecht</strong> (NL) &#8212; van den Berg M et al. 2023, <em>Computers in Biology and Medicine</em>. <a href="https://doi.org/10.1016/j.compbiomed.2023.107156">DOI: 10.1016/j.compbiomed.2023.107156</a>. Largest single-center cohort; alarm-fatigue framework.</p></li><li><p><strong>Eindhoven / M&#225;xima</strong> (NL) &#8212; Yang M et al. 2024, <em>Computer Methods and Programs in Biomedicine</em>. <a href="https://doi.org/10.1016/j.cmpb.2024.108335">DOI: 10.1016/j.cmpb.2024.108335</a>. Most thorough feature ablation; Sino-Dutch collaboration.</p></li><li><p><strong>Antwerp / Innocens</strong> (BE) &#8212; Meeus M et al. 2024, <em>Journal of Pediatrics</em>. <a href="https://doi.org/10.1016/j.jpeds.2023.113869">DOI: 10.1016/j.jpeds.2023.113869</a>. Joint LOS+NEC prediction; only commercial spinoff.</p></li><li><p><strong>Karolinska / KTH</strong> (SE) &#8212; Honor&#233; A et al. 2023, <em>Acta Paediatrica</em>. <a href="https://doi.org/10.1111/apa.16660">DOI: 10.1111/apa.16660</a>. Bridges both branches; senior co-authors wrote the Persad 2021 SR.</p></li><li><p><strong>Rennes</strong> (FR) &#8212; Leon C et al. 2021, <em>IEEE Journal of Biomedical and Health Informatics</em>. <a href="https://doi.org/10.1109/JBHI.2020.3021662">DOI: 10.1109/JBHI.2020.3021662</a>. Visibility-graph HRV; methodologically distinctive.</p></li><li><p><strong>Lausanne</strong> (CH) &#8212; Rio L et al. 2022, <em>Pediatric Research</em>. <a href="https://doi.org/10.1038/s41390-021-01913-9">DOI: 10.1038/s41390-021-01913-9</a>. Only independent HeRO validation outside the US.</p></li></ul><div><hr></div><p>Two synthesis papers anchor the field: Koppens 2023 (Branch A) and Kainth 2024 (Branch B).</p><p>A few observations on this map.</p><p><strong>Geography.</strong> Four of the seven groups are in continental Western Europe, two Dutch centres, plus Belgium, France, Switzerland, and one is in Sweden if you count the Karolinska bridge group. One is in the US. <strong>Couldn&#8217;t find any in Asia, Africa, or Latin America.</strong> I checked specifically for Chinese Branch A work, because the Chinese neonatology literature is data-rich and methodologically capable. The Chinese research bet in this space is firmly on Branch B, lab-feature and clinical-variable nomograms. The continuous-physiology ML problem is, almost entirely, a Western European and American research phenomenon. </p><p><strong>Time.</strong> The field has gone a bit quiet in 2024&#8211;2026. I checked the 184 most recent papers after the Kainth SR cutoff. Almost none are new Branch A LOS prediction work from the major groups. The work has broadened, to NEC, to catheter-related bloodstream infections, to LMIC adaptation, to regulatory questions, but it has not deepened. There has been no major new model paper from any of the seven groups in roughly 18 months.</p><h2>What every group is finding</h2><p>This is the part I expected to be messy. It turned out to be cleaner than I thought.</p><p>Across all seven anchor papers, different signal stacks (1-channel to 7-channel), different sampling rates (1/h to 250 Hz), different ML methods (logistic regression to deep learning), different cohort sizes (51 to 3,151), <strong>AUCs converge on 0.78&#8211;0.88 at the moment of clinical suspicion.</strong> The single external validation in the field (Kausch trained on UVA, tested on Columbia and St Louis) drops AUC by about 0.03 across centres. The convergence is real, and it isn&#8217;t a methodological artifact.</p><p>The more clinically meaningful number is <strong>what fraction of LOS episodes are detected before clinical recognition:</strong></p><ul><li><p><strong>Berg 2023:</strong> 47&#8211;60% of patients detected pre-clinically, using a multi-threshold alarm policy.</p></li><li><p><strong>Kausch 2023:</strong> significant risk elevation 23&#8211;24 hours before blood culture.</p></li><li><p><strong>Yang 2024:</strong> 96% patient-wise detection before clinical suspicion, but on a small cohort (n=119).</p></li><li><p><strong>Meeus 2023:</strong> 69% of all episodes, 81% of severe episodes; median time gain 10 hours.</p></li></ul><p>The consensus is this: <strong>roughly half to two-thirds of LOS episodes appear to be detectable hours before clinical recognition with current technology.</strong> That is the technical answer to the central question of the field.</p><p>Two methodological points that matter more than they get credit for:</p><blockquote><p><strong>Yang&#8217;s sampling-rate ablation is the most underrated paper in this list.</strong> They compared raw waveforms (AUC 0.886) &#8594; 1 Hz vitals (0.875) &#8594; 1/min vitals (0.825) &#8594; 1/h vitals (0.687). The 1/min-to-1/h jump destroys the signal. The 1 Hz-to-1/min jump costs about 5 AUC points. What this tells you is that <strong>minute-by-minute data works fine</strong>; once you summarize to hourly aggregates, the signal collapses. That&#8217;s the floor of deployable signal resolution, and it&#8217;s empirically settled now.</p></blockquote><blockquote><p><strong>Kausch&#8217;s pulse-rate-equals-ECG finding is the most underrated result.</strong> They showed that POWS performance was within 0.01 AUC whether the heart-rate signal came from ECG or from pulse oximetry. That means a Branch A model can run on a standalone pulse oximeter, no ECG leads, no specialized monitor. For low-resource deployment, that is the difference between a research curiosity and a deployable tool.</p></blockquote><h2>What every group is choosing not to do</h2><p>A consistent finding across the papers that compared multiple ML methods: <strong>none of them found that complex models beat simple ones.</strong> Berg compared logistic regression, GAMs, and XGBoost, all within 0.01 AUC, chose LR for interpretability. Kausch compared LR with cubic splines, neural networks, XGBoost, and random forest, all within 0.01 AUC, chose LR (&#8221;equal performance, better explainability&#8221;). Yang compared seven methods including two deep-learning architectures, gradient boosting won; deep learning actually underperformed on their cohort.</p><p>This is unusual for an ML field in 2024&#8211;2026. It means LOS prediction is <strong>not a deep-learning-shaped problem (yet?).</strong> It is a feature-engineering problem with a clear physiological signal stack, and once you have the right features, linear and tree-based methods reach the same ceiling. The interesting work is in choosing features, not architectures.</p><p>Said differently: <strong>the bottleneck in this field is not algorithm sophistication. It&#8217;s the things ML alone can&#8217;t fix.</strong></p><h2>Where I think the WKZ contributed</h2><p>I&#8217;ll claim one specific thing for our group, because reading the comparative literature makes it visible in a way it wasn&#8217;t when we published.</p><p>The Berg 2023 paper proposed a specific framework for evaluating these models the way a clinician would actually experience them: <strong>continuous hourly prediction, multi-threshold alarms with an 8-hour refractory period (the length of a NICU shift), and explicit accounting for false-alarm burden as alarms per patient-day.</strong> The point was to stop optimizing for AUC alone and start optimizing for what bedside clinicians would tolerate.</p><p>Within 12 months, two other groups had adopted the same framework. Yang 2024 cites Berg explicitly as the basis for their multi-threshold alarm policy. Meeus 2023 echoes the alarm-rate accounting. The 8-hour shift-length refractory period and the threshold-escalation rule are now a small but real convention in this corner of the field.</p><p>I&#8217;m not claiming we solved LOS prediction. The actual modelling work is roughly equivalent across the major groups. What I&#8217;ll claim is that <strong>WKZ set a convention for how to evaluate these models for deployment</strong>, and the field adopted it. That&#8217;s worth saying.</p><h2>The honest gap: the deployment vacuum</h2><p>Here is what is <em>not</em> in the literature, despite 15 years of work and seven productive academic groups:</p><ul><li><p><strong>Zero prospective validation studies</strong> of any of these models in current clinical use.</p></li><li><p><strong>One randomized controlled trial.</strong> The <a href="https://doi.org/10.1016/j.jpeds.2011.06.044">Moorman 2011 HeRO trial</a>. Fourteen years old. Mortality benefit ARR 2.07%, NNTM 49 (95% CI 24&#8211;15,484). High risk of bias per the Cochrane assessment. The long-term follow-up (King 2021) showed persistent mortality reduction at 18&#8211;22 months, but also an unexplained increase in deafness in the HRC-monitored arm (4.4% vs 0.5%), possibly aminoglycoside-related.</p></li><li><p><strong>One independent external validation</strong> of the only commercial product. Rio 2021, single-centre Lausanne data, showed real-world HeRO performance is strongly gestational-age-dependent: sensitivity 76% in infants below 28 weeks, falling to 25% above 32 weeks. The optimal threshold in their cohort was 2.76, not the FDA-cleared 2.0.</p></li><li><p><strong>One commercial deployment</strong> outside HeRO: Innocens BV, the Antwerp spinoff. Currently navigating the European MDR/CE certification pathway. <a href="https://doi.org/10.2196/72809">Vrijlandt et al. 2026</a> (Erasmus MC) is the first published paper to think through what that regulatory pathway looks like for paediatric clinical decision support in Europe.</p></li></ul><p>The Koppens 2023 SR lands the policy verdict most directly: <em>&#8220;methodological weaknesses and limited generalizability do not justify implementation of HRC in clinical care. A large international RCT is warranted.&#8221;</em></p><p>That is the gap. We have seven academic groups, four major model papers, two systematic reviews, one underpowered 2011 RCT, one independent external validation showing real-world drift, and one commercial product navigating European regulation. <strong>The technical question &#8220;can we predict LOS pre-clinically with ML?&#8221;  has been answered, &#8220;reasonably well&#8221;. The deployment question is wide open.</strong></p><h2>What this means for CEPAS</h2><p>I came into this search expecting to summarize the technical state of the art. I&#8217;m leaving it convinced that&#8217;s the wrong talk to give.</p><p>The technical state of the art is converged. AUCs at 0.78&#8211;0.88. Half to two-thirds of LOS episodes detectable hours before clinical recognition. Feature engineering, not architecture, drives the ceiling. Pulse-rate-only signal stacks work. Alarm-fatigue frameworks are now field convention.</p><p>What is missing is everything that turns a model into a deployable clinical tool: prospective trials, international validation, regulatory clarity, deployment infrastructure outside high-income academic centres, and an honest accounting of what real-world performance looks like once you leave your training cohort.</p><div><hr></div><p><strong>Episode 3</strong> will go deeper on the head-to-head between the major model papers, specifically on signal stack choices and alarm policy design, because those decisions matter more than the ML methodology and almost nobody writes about them clearly. After that, the series turns toward deployment and the European regulatory frame.</p><p><strong>If you spotted a group I missed in the map, or have a paper I should have anchored on, please tell me. The whole point of doing this transparently is to fix what I get wrong. </strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/p/episode-2-mapping-the-field/comments&quot;,&quot;text&quot;:&quot;Laat een reactie achter&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://road2cepas.substack.com/p/episode-2-mapping-the-field/comments"><span>Laat een reactie achter</span></a></p><div><hr></div><p><em>This post is part of &#8220;Road to CEPAS 2026&#8221;, a transparent research diary as I prepare a 20-minute talk for the CEPAS 2026 congress in Lyon. <a href="https://road2cepas.substack.com/p/the-honest-problem">Episode 1</a> traced the HeRO/Moorman lineage; this episode maps the current state of the field. The full chat transcript that produced this post will be available on <a href="https://neonatology.ai/">neonatology.ai</a>.</em></p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Abonneer nu&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://road2cepas.substack.com/subscribe?"><span>Abonneer nu</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Honest Problem]]></title><description><![CDATA[Road to CEPAS 2026 &#8212; Episode 1]]></description><link>https://road2cepas.substack.com/p/the-honest-problem</link><guid isPermaLink="false">https://road2cepas.substack.com/p/the-honest-problem</guid><dc:creator><![CDATA[Quantum Neonatology]]></dc:creator><pubDate>Wed, 13 May 2026 18:01:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oP-b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In late October 2026 I&#8217;ll stand on a stage in Lyon for twenty minutes and talk about big data, AI, and sepsis prediction in the NICU. The session is called <em>Omics in Sepsis</em>. The venue is the first <a href="https://www.cepas.org">Congress of the European Paediatric Academic Societies</a>. I said yes the moment the invitation arrived.</p><p>I&#8217;ve been quietly uneasy about it ever since.</p><p>Not about the speaking part, that&#8217;s familiar territory. The unease is about what I&#8217;m supposed to <em>say</em>. Because if I&#8217;m honest with the audience, and I intend to be, the talk has to begin with a problem most of us in this field have stopped acknowledging out loud.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oP-b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oP-b!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic 424w, https://substackcdn.com/image/fetch/$s_!oP-b!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic 848w, https://substackcdn.com/image/fetch/$s_!oP-b!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic 1272w, https://substackcdn.com/image/fetch/$s_!oP-b!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oP-b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:337280,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/heic&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://road2cepas.substack.com/i/196910333?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oP-b!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic 424w, https://substackcdn.com/image/fetch/$s_!oP-b!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic 848w, https://substackcdn.com/image/fetch/$s_!oP-b!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic 1272w, https://substackcdn.com/image/fetch/$s_!oP-b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0041a64-02ea-4816-966b-b181012759eb_1536x1024.heic 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>ChatGPT Image 2</em></p><h2>We can predict. So what?</h2><p>We can predict late-onset sepsis in preterm neonates reasonably well, and we have been able to for years. The literature is full of models. ROC curves climbing toward 0.80, 0.85, occasionally higher with enough feature engineering and a cooperative dataset. That includes &#8212; full disclosure &#8212; our own paper from 2023 in <em><a href="https://www-sciencedirect-com.utrechtuniversity.idm.oclc.org/science/article/pii/S0010482523006212?via%3Dihub">Computers in Biology and Medicine</a></em>. AUC 0.73 at the moment of clinical suspicion. A longitudinal impact simulation suggesting we could have flagged 47% of LOS cases ahead of the bedside team, while keeping false alarms below three per day. A genuinely useful exercise in clinical impact assessment, more than most papers in the space at the time.</p><p>It&#8217;s a good paper. I&#8217;m proud of the work. It is also a model that, as I write this in the spring of 2026, is not being used to take care of a single baby. Not in our NICU. Not anywhere.</p><p>That is the honest problem. Not <em>can AI predict neonatal sepsis</em> <em>or at least clinical deterioration,</em> that question is largely settled. The honest problem is the next one. <strong>Once we have the prediction, what do we actually do with it?</strong></p><h2>The HeRO question</h2><p>Consider the <a href="https://www.heroscore.com">HeRO monitor</a> &#8212; heart rate characteristics, FDA-cleared, with a randomized trial showing absolute mortality reduction in very low birthweight infants. That trial was published in 2011. Fifteen years on, HeRO is on the market, the evidence is real, and yet ask any group of neonatologists how they respond when the score climbs and you&#8217;ll get a different answer at every cot. Order an IL-6? A CRP? Both? Neither? Increase observation? Start empiric antibiotics? Wait? Blood culture?</p><p>This is not a HeRO problem. It is the problem of every prediction tool we have built, including ours. We have invested decades into the <em>predict</em> half of the equation and almost nothing into the <em>act</em> half. There is no validated response protocol. No agreed-upon downstream pathway. No randomized comparison of &#8220;alarm fires &#8594; action A&#8221; versus &#8220;alarm fires &#8594; action B.&#8221; The score lights up and the clinical reasoning that follows is, essentially, vibes.</p><p>That is the gap I want to talk about in Lyon. Not the success stories. The gap.</p><h2>What this project is</h2><p>So, between now and 31 October 2026, I&#8217;m going to build the talk in public. Twenty-some weeks. Posts on this Substack, a growing public literature base, an AI-assisted mini systematic review along the way, and the slides themselves at the end. Everything sourced. Everything visible. Nothing dressed up.</p><p>I&#8217;m doing it with Claude as a research partner. Not as a gimmick, as the actual working method. PubMed searches, evidence tables, critical re-reads of papers (including my own), regulatory landscape mapping, slide logic. Every post will include a short &#8220;How I used Claude&#8221; section so you can see exactly where the AI helped and where it didn&#8217;t. If you want to copy the workflow for your own project, you&#8217;ll have everything you need.</p><p>The thesis I&#8217;ll keep returning to is simple. <strong>Prediction is not clinical utility, and the field has not yet built the bridge between them.</strong> That bridge is regulatory, methodological, behavioural, and economic. Each of the next posts is a piece of it.</p><h2>How I used Claude in this post</h2><p>I drafted the framing and the argument. Claude turned the bullet points of the project brief, the CEPAS invitation, the paper details, the &#8220;dal&#8221; I find myself in, into structured prose, kept the word count in range, and pushed back on a softer opening I&#8217;d originally written. I edited every paragraph. The unease is mine.</p><h2>What&#8217;s next</h2><p>Episode 2 lands soon: <em>Searching the Evidence &#8212; Live.</em> I&#8217;ll run a real PubMed search for LOS prediction models with Claude using the PubMed MCP, on screen, queries and all. If you want to follow along or copy the method, subscribe, and meet me in Lyon, virtually or otherwise.</p><p>More will appear soon @ <a href="https://neonatology.ai/">neonatology.ai</a>.</p><p>NB the images will become more happy along the way </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://road2cepas.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Abonneren&quot;,&quot;language&quot;:&quot;nl&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Substack van Quantum! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Typ je e-mailadres&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Abonneren"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item></channel></rss>