A wave of AI-native platforms is pitching itself directly at the design controls and regulatory affairs function: feed the tool your design inputs, your risk data, your dev-tool activity, and it hands back drafted documentation, a live traceability matrix, and a risk or compliance score. The pitch is seductive because the pain it targets is real. Design History File compilation, traceability matrix maintenance, and submission prep are genuinely slow, genuinely manual, and genuinely error-prone when done by hand across a growing document set. I want to take a skeptical, technical look at three platforms working this space right now: MedHive.ai, Ketryx, and QualiVerse from Quality Means Business (QMB Inc.). What each claims, and what I think still needs a human regulatory reviewer sitting on top of it.
MedHive.ai
MedHive positions itself as an AI-based knowledge management platform for medical device R&D, built on what it calls a neuro-symbolic AI approach, meaning it pairs a knowledge graph and rule-based reasoning with a language model rather than relying on the language model alone. The platform's stated features include automated artifact classification, automated artifact generation, automated change management, and per-artifact accuracy scoring, plus a 510(k) drafting capability that the company says has saved individual teams weeks of premarket notification preparation time. MedHive markets itself against pure LLM tools, generic PLM/QMS platforms like Siemens Teamcenter and PTC Windchill, and QMS-adjacent tools like Greenlight Guru and Qualio, arguing that the combination of domain-specific context and a persistent knowledge graph produces more reliable output than a general-purpose chat interface pointed at your documents.
The neuro-symbolic framing and the per-artifact accuracy scoring are worth taking seriously on their technical merits: a knowledge graph that tracks relationships between design inputs, outputs, and historical decisions is a more defensible architecture for a document-generation tool than pure retrieval-augmented generation over a document dump, because it constrains the model's output against an explicit structure rather than hoping the right context gets retrieved. That said, MedHive's public marketing leans heavily on aggregate efficiency claims, faster release cycles, reduced documentation time, hours saved per developer, without publishing the methodology behind those figures or third-party validation of them. Reasonable to expect from an early-stage platform's marketing site, but worth treating as a starting point for due diligence, not a substitute for it.
Ketryx
Ketryx is the most mature and most aggressively positioned of the three, with a public case study roster that includes Meta Reality Labs, Flo Health, DeepHealth, HeartFlow, Cytovale, and Beacon Biosignals, and a claim that four of the top five Fortune 500 MedTech companies now run on the platform. The core architecture connects to a team's existing dev tools, Jira, GitHub, GitLab, AWS, and layers AI agents on top that draft documentation, propose traceability links, flag QMS process deviations in real time, and automatically compile what Ketryx calls a submission-ready Design and Development File. The company's headline efficiency claim is a documentation-time reduction of up to 90%, with individual case studies citing figures like a 70% faster change-impact assessment and a 60% reduction in documentation cycle time.
Technically, Ketryx's strongest claim is traceability automation: continuously mapping requirements, risks, tests, and code across connected systems, and having AI agents detect coverage gaps and propose missing links before an auditor finds them. That's a genuinely different capability from document drafting. A traceability matrix that updates itself as engineers work in Jira and commit code in GitHub, rather than requiring someone to manually reconcile spreadsheets after the fact, addresses one of the most persistent and error-prone parts of design controls, and it's now an explicit regulatory expectation under the QMSR's incorporation of ISO 13485 Clause 7.3, not just a best practice. I'd flag two limits on that claim, though. First, the quality of an automated traceability matrix is bounded by the quality of what's captured in the connected systems; if an engineer's Jira ticket doesn't actually describe the requirement correctly, the tool will faithfully trace a wrong thing with full confidence. Second, Ketryx's public case studies are heavily weighted toward software and SaMD compliance, IEC 62304, cybersecurity, AI/ML validation, where the artifacts genuinely live inside connected dev tools. The track record for hardware-heavy device development, where a meaningful share of design control work involves mechanical test data, biocompatibility studies, and sterilization validation that don't naturally flow through Jira or GitHub, is thinner in the public case study set. That's not necessarily a capability gap in the platform itself, since Ketryx does support hardware-software products as a stated use case, but it means the "reduce documentation time by 90%" figure is best read as most reliably demonstrated for software-centric device work, not as a blanket claim across every design control discipline.
QualiVerse (QMB Inc.)
QualiVerse, built by Quality Means Business, is the most submission-focused of the three, framed explicitly around FDA premarket work: product code classification, predicate device identification, and simulated submission review, backed by what the company calls a proprietary 7-point AI risk model. Public claims include a 95% submission success rate, 80% faster submission prep time, and the ability to generate "thousands of pages" of submission-ready documentation benchmarked against real-world FDA data. The company raised a $2 million seed round in November 2025 to formally launch the platform and is the youngest and least established of the three by a meaningful margin, both in company age and in the depth of its public case study material, which currently consists of a single named customer testimonial.
The predicate-matching and product-code-classification capability is the piece I'd scrutinize hardest. Identifying candidate predicate devices and product codes is fundamentally a search and pattern-matching task against FDA's own 510(k) database and product classification system, which makes it a genuinely good fit for AI acceleration; a tool that surfaces the right candidate predicates faster than manual database searching is real time saved. But predicate selection isn't just data retrieval. It's a regulatory strategy decision that has to hold up as a substantial equivalence argument, one that accounts for intended use, technological characteristics, and the specific comparative claims FDA will scrutinize for that device type. A "95% submission success" figure, without a published definition of what counts as success (RTA acceptance versus first-cycle clearance versus eventual clearance after additional information requests) or a disclosed sample size, is the kind of marketing metric I'd want unpacked in a sales conversation before treating it as evidence the AI is making sound regulatory strategy calls rather than good first-pass drafts that a reviewer then has to substantially rework.
The pattern across all three
None of these platforms claims that its AI clears a device or replaces a regulatory reviewer's judgment. That's worth noting explicitly, because it's the tell that the vendors themselves know where the line sits, even when their marketing copy leans hard on speed and automation. Every one of them frames itself as "expert-in-the-loop," "AI Agents that surface changes for human review," or an accelerant to a human-led process, not a replacement for one. That framing lines up with what I'd expect technically. Drafting a document from a template and known inputs, building and maintaining a traceability matrix across connected systems, and surfacing candidate predicates from a structured database are all pattern-matching and structuring tasks AI is well suited for. Assigning severity and occurrence scores in a risk analysis, judging whether a substantial equivalence argument will actually hold up against FDA's specific concerns for a device type, and deciding whether an automatically detected traceability gap represents real risk or a false positive from an incompletely configured system are judgment tasks that require an engineer or regulatory professional who understands the specific device, not just the pattern.
The risk I'd flag to anyone evaluating these tools isn't that the AI gets things wrong outright. It's the more subtle failure mode: an automatically generated document or automatically maintained traceability matrix that looks complete and well-structured, because the tool is genuinely good at producing well-structured output, but that inherits an error or a gap from upstream data that nobody caught because the output looked too finished to warrant a hard second look. That's the same failure mode I described in the previous piece on AI-assisted design FMEA population, and it holds here too. These platforms compress the authoring work. They don't compress the review work, and treating the two as equivalent is where the actual regulatory risk lives.
medhive.ai, ketryx.com/product, ketryx.com/capabilities/traceability, qmb.ai, Quality Means Business (QMB Inc.) Raises $2M Seed to Launch QualiVerse®, BusinessWire, Nov. 12, 2025. All platform capability and metric descriptions in this piece are drawn from each company's own public marketing material as of August 2026; none of the efficiency or success-rate figures cited by the vendors have been independently verified by the author.
