Does DITA XML Reduce AI Hallucinations?Correctly structuring content as a DITA task doesn’t establish whether that content includes the information needed to answer a questionFormer clients have asked,
They were caught off guard when I answered, “Not necessarily.” I followed that declaration with this statement: Using DITA XML to create documentation doesn’t, on its own, mean an AI answer engine will provide more accurate answers.There was quiet in the meeting room as the team considered my words. One writer spoke up, breaking the silence to ask,
”You’re right,” I replied, “but only if (and when) the system can retrieve that information, and subsequently, recall it and act upon it. Tech writers must produce accurate facts and clearly explain the relationships customers (humans and machines) need to understand the instructions. AI systems must find the right information and apply it correctly.” Producing valid DITA topics doesn’t guarantee complete content. Consider, for example, a document that follows DITA rules, but fails to state an important prerequisite or another dependency. When that happens, an AI answer engine may be forced to guess (infer) the missing information and perhaps get it wrong. My first experience with XML authoring came in 1999 at a pharmaceutical company trying to improve how it produced new drug applications. These submissions could approach 100,000 pages and contained information the U.S. Food and Drug Administration needed to evaluate drugs for approval. The company applied information science to producing and managing that content as reusable XML components. I’ve advocated for semantic structured content ever since. Much of my subsequent consulting work involved helping tech docs teams adopt DITA, an XML specification, and a component content management system (CCMS). What DITA IdentifiesTech writers use DITA markup to create content in accordance with the DITA specification. A DITA’s Can Valid DITA XML Leave An Answer Unstated?Yes, it can. And, perhaps more often than some might realize. The docs we produce can conform to the DITA specification while at the same time omitting facts or relationships needed to answer a customer’s question. A correctly structured restoration task, for example, might never state how long trial account users have to restore deleted projects. Validity also doesn’t establish how fully a team uses DITA’s capabilities. Authors may use less specific markup than the specification allows or omit metadata that would clarify when instructions apply. If the required facts or relationships aren’t explicit elsewhere in the information the engine receives, it may infer them incorrectly. More detailed markup can supply useful cues, but the appropriate level of detail depends on the content and the questions it needs to support. Suppose an answer engine receives two DITA topics, including their relevant markup and metadata. The topics contain these statements:
The customer’s question is:
A product could retain deleted project data for 30 days while giving users 14 days to restore it. If no other available information establishes the restoration period for trial accounts, an answer of “30 days” depends on assuming that retention and restoration have the same duration. An engine could make that assumption incorrectly or recognize that the source |