A client asked me to review a device development grant proposal before it went out the door to a funder. The device was a continuous bioelectrical impedance spectroscopy system, a non-invasive sensor platform measuring multi-frequency electrical impedance through the body to track fluid-related physiological changes, aimed at continuous fluid status monitoring in heart failure, renal disease, and perioperative care. The proposal benchmarked itself against a commercially available predicate, the ImpediMed SFB7, and framed a future 510(k) submission around that comparison. On paper it read well. Underneath it, the design controls foundation had gaps that would matter a great deal the moment this moved from a grant narrative into an actual regulatory submission, and some of them were the kind that are easy for a first pass, human or AI, to miss.
What tipped off the review
Milestone 1 of the proposal specified a set of finished performance numbers with no visible derivation behind them: a 5 to 500 kHz frequency range, measurement error at or below 5 percent, coefficient of variation under 5 percent, drift under 2 percent over 30 minutes, all stated as benchmarks against the predicate device. Those numbers read like design requirements. They aren't. Under a design-controls-compliant process, you'd expect to find three linked things behind a spec like that: a User Needs statement describing the clinical decision the data has to support, a Design Input Requirements document translating that need into a measurable specification with a stated rationale, and a traceability reference tying each input to the verification activity that confirms it. None of that existed in the proposal. What existed instead was the verification test's own pass/fail threshold, presented as if it were the requirement that generated the test, rather than the test that confirms the requirement.
That's a subtle but important inversion, and it's worth being precise about why it matters. The most likely explanation, and this is an inference rather than a confirmed fact, is that the numbers were set to whatever the predicate device already achieves. That's a reasonable competitive benchmark for a proposal narrative. It is not the same thing as a documented, clinically justified design input, and a reviewer or auditor who asks "why 5 percent and not 3 percent or 8 percent" deserves an answer grounded in clinical need, not "because that's what the predicate does." A corrected version of this section wouldn't be complicated to build: a short User Needs statement naming the clinical decision the data needs to support, something like clinicians need continuous fluid trend data sufficient to inform a treatment decision, a Design Input table with each spec, accuracy, CV, drift, frequency range, tied back to that need with an actual rationale, and a traceability line showing which verification activity in Milestone 1 confirms which input. Three linked documents instead of one table of numbers with no lineage behind them.
What a first-pass AI review caught on its own
I ran an initial review pass with an AI model against the proposal text, and it's worth being honest about how much it caught without much prompting. It flagged the invisible design inputs described above. It flagged that no risk management file existed in the proposal at all, and that risk analysis was scoped as a single end-of-project deliverable rather than the live, updated-throughout-development process ISO 14971 actually describes. It caught that there was no software safety classification anywhere for the signal-processing algorithm, despite that algorithm driving the entire clinical output of the device, which is a meaningful gap given how central software classification is to scoping the right level of software lifecycle rigor under IEC 62304. It caught a configuration-control problem: a hardware platform transition was scheduled to overlap with the benchmarking data collection period, which raises an obvious question about whether the data being collected will actually represent the device that eventually gets submitted. And it flagged a mislabeling problem worth dwelling on: a mechanistic verification study, healthy volunteers under a controlled fluid load, was described in the narrative sections in language that made it sound like design validation in the intended patient population. Under ISO 13485 Clause 7.3.7, those are not the same activity, and treating one as if it satisfies the other is exactly the kind of category error that creates real problems downstream, when an actual validation study still needs to happen and nobody budgeted time or funding for it because the proposal implied it was already done.
That's a genuinely useful first pass. It's also, notably, entirely a document-analysis exercise: reading defined terms against a defined framework and flagging where the document doesn't match the framework's expectations.
Where the human judgment actually happened
Four things in this review needed a person who has actually run a design control program, not just a model that can pattern-match a proposal against ISO 13485 and IEC 62304.
The first is calibration. An AI can tell you a document lacks a formal risk management file. It can't tell you, from experience, how often early-stage teams have partial versions of this work sitting somewhere that never made it into a grant narrative because nobody thought a funding proposal needed a risk management appendix. That distinction matters for how you phrase the finding. "This proposal has no documented risk management process" is accurate and also, without qualification, reads as an accusation that the team hasn't done any risk thinking at all, which may or may not be true and isn't something the document itself can tell you. Phrasing that gap as a documentation finding rather than a competence finding is a judgment call, and it's the kind of judgment call that determines whether the client's relationship with their own team survives the review.
The second is cross-referencing sections the way a single-pass analysis doesn't naturally do. The proposal described its clinical study as passive monitoring in nearly every section. Buried in one budget line item, in language that had nothing to do with the clinical description elsewhere, was a reference that made clear the actual protocol involved an interventional IV fluid-loading procedure, not passive observation. That's not a subtle contradiction if you're reading the budget narrative and the clinical narrative as one connected document and specifically looking for places where they might disagree. It's exactly the kind of thing a document-by-document or section-by-section pass treats as independent facts rather than a discrepancy, and the distinction is regulatory-classification-relevant: an interventional protocol carries different risk categorization and different informed consent obligations than passive monitoring does. Catching it required deliberately reading across sections for consistency, not just scoring each section against a checklist.
The third is regulatory strategy judgment, which is a different skill from document review entirely. The substantial equivalence argument underpinning the whole 510(k) plan was, on its own terms, a strategic risk worth flagging directly: continuous monitoring against a predicate that performs single-timepoint measurement is precisely the kind of technological characteristic difference that can defeat a substantial equivalence claim, or at minimum draw a request for additional data that the current milestone plan didn't budget time for. Recognizing that requires understanding how FDA actually evaluates SE arguments, not just confirming that a predicate was named somewhere in the document.
The fourth is audience and tone, and it's worth naming explicitly because it's easy to treat as a soft skill rather than a technical one. This document was headed to a funder, not an internal file. A first-pass AI review can hand you a technically accurate list of every gap in the proposal. It can't tell you how to say it to the audience the document is actually written for, without either overstating how bad the gaps are or underselling the real risk sitting underneath a polished narrative. Reading the proposal for language that would read as an unearned claim to a sharp funder reviewer, "validated" where the honest word was "verified in a mechanistic model," "risk-managed" where the honest phrase was "risk considerations acknowledged," is a judgment call about how the document will actually be read by the person it's for. That's not something you outsource, no matter how good the underlying gap analysis is.
The honest takeaway
The AI review here did real, substantive work fast: it read a document against two regulatory frameworks and surfaced five distinct, correct findings without much steering. What it couldn't do, and what took the bulk of the actual judgment in this engagement, was calibrate how to say what it found, catch a contradiction that only exists across two sections read together rather than within either one, and make a genuine regulatory strategy call about whether a predicate comparison would hold up. Those aren't limitations that better prompting fixes. They're the difference between analyzing a document and understanding what the document is actually for, who's going to read it, and what happens to the client if the gaps get found later instead of now.
