Quick Answer
In eCTD 3.2.2, index.xml organizes the submission and references the regional Module 1 backbone. A leaf connects a document's title, location, identity, checksum and lifecycle instruction. XML parsing alone cannot establish that those relationships are correct. FDA eCTD v4.0 uses a different submissionunit.xml message model, so do not carry the older backbone template across unchanged.
An XML file can open successfully while pointing to the wrong document. A checksum can match while the document belongs to the wrong section. A technically correct reference can still carry an inappropriate lifecycle instruction. Those are different failures, and diagnosing them separately makes a publisher's validation report easier to investigate.
This tutorial teaches that separation through a complete local exercise. You will create a small fictional XML record and payload, inspect them, introduce three faults, and restore the working example. You do not need a publisher license or access to an agency gateway.
Scope: The eCTD concepts discussed here concern the 3.2.2 backbone. The supplied practice.xml deliberately uses a teaching wrapper and text payload; it is neither an eCTD sequence nor a submission template. No result below means that a dossier is technically valid or accepted by an authority.
Locate the backbone before investigating a leaf
For the 3.2.2 model, ICH places index.xml and its index-md5.txt checksum file in the sequence directory. The regional XML belongs in the regional Module 1 directory, rather than beside index.xml. The main backbone references that regional file. Its supporting DTD files sit within the utilities structure. See the ICH eCTD 3.2.2 specification, Appendix 6.
For orientation, the following deliberately incomplete map shows relationships, not every required submission component:
application/
0000/
index.xml
index-md5.txt
m1/
us/
us-regional.xml
m2/
...documents...
util/
...applicable DTD and presentation files...The dots are labels, not files to create. Region-specific filenames, applicable specifications, and actual content must come from the selected regional implementation. Use the Module 1 preparation guide to establish that regional scope, and the module structure reference to place the content in context.
Read the attributes as separate questions
| Information | Question to ask while investigating |
|---|---|
| Leaf ID | Which XML record are we talking about? |
| Title | What will the reviewer understand this document to be? |
| File reference | Which file does the record point to, from this XML location? |
| Checksum | Does the referenced file's byte content match the recorded digest? |
| Operation | What change is intended in the submission history? |
| Modified-file reference, where applicable | Which earlier leaf is affected? |
These questions are an inspection aid. They do not replace the governing DTD or regional validation rules. In particular, “the file exists” answers neither whether it is the approved file nor whether the chosen operation is appropriate.
Prepare the practice environment
Use Python 3.9 or newer with the standard library and a terminal. Confirm that python3 --version returns the interpreter you intend to use; on some Windows installations the equivalent command is py -3. No packages or credentials are required.
Save the complete code below as inspect_backbone.py in an ordinary working directory. Run only this supplied fixture. It creates a uniquely named temporary directory, performs its checks, and removes that directory automatically. It does not open a real dossier, contact a service, or modify an existing application.
The exercise supplies one payload containing the exact bytes Practice document version one followed by a newline. A fictional leaf called practice-001 points to that file. Its checksum is calculated from those bytes at setup, so there is no unexplained digest to copy from a screenshot.
The namespace string in this example follows the historical spelling in the ICH 3.2.2 DTD. Treat namespace URIs as identifiers: changing one because a different spelling looks more familiar changes what a parser sees. The wrapper remains explicitly nonstandard, and no DTD is declared or validated.
Run the complete inspection exercise
Copy the whole block, including the imports and final with block:
from copy import deepcopy
from hashlib import md5
from pathlib import Path, PurePosixPath
from tempfile import TemporaryDirectory
import xml.etree.ElementTree as ET
XLINK = "http://www.w3c.org/1999/xlink"
HREF = "{" + XLINK + "}href"
PAYLOAD = b"Practice document version one\n"
def digest(content):
# Protocol-style byte comparison, not a security guarantee.
return md5(content, usedforsecurity=False).hexdigest()
def inspect(xml_bytes, folder):
try:
root = ET.fromstring(xml_bytes)
except ET.ParseError:
return "MALFORMED_XML"
leaves = root.findall("leaf")
if root.tag != "practice" or len(leaves) != 1:
return "UNEXPECTED_PRACTICE_STRUCTURE"
leaf = leaves[0]
if leaf.get("ID") != "practice-001":
return "UNEXPECTED_PRACTICE_ID"
if leaf.get("operation") != "new":
return "OUTSIDE_PRACTICE_OPERATION"
if leaf.get("checksum-type") != "md5":
return "OUTSIDE_PRACTICE_CHECKSUM"
if not (leaf.findtext("title") or "").strip():
return "MISSING_TITLE"
reference = leaf.get(HREF, "")
# This lesson accepts one literal relative POSIX path only.
if not reference or any(c in reference for c in ("\\", ":", "?", "#", "%")):
return "OUTSIDE_PRACTICE_PATH"
relative = PurePosixPath(reference)
if relative.is_absolute() or ".." in relative.parts:
return "OUTSIDE_PRACTICE_PATH"
base = folder.resolve()
target = (base / reference).resolve()
if not target.is_relative_to(base):
return "OUTSIDE_PRACTICE_PATH"
if not target.is_file():
return "MISSING_FILE"
if digest(target.read_bytes()) != leaf.get("checksum"):
return "CHECKSUM_MISMATCH"
return "PRACTICE_CHECKS_PASS"
def checkpoint(label, actual, expected):
print(label + ": " + actual)
if actual != expected:
raise RuntimeError(label + ": expected " + expected)
with TemporaryDirectory(prefix="assyro-xml-practice-") as location:
folder = Path(location)
payload_path = folder / "documents" / "practice.txt"
payload_path.parent.mkdir()
payload_path.write_bytes(PAYLOAD)
root = ET.Element("practice")
leaf = ET.SubElement(root, "leaf", {
"ID": "practice-001",
"operation": "new",
"checksum-type": "md5",
"checksum": digest(PAYLOAD),
HREF: "documents/practice.txt",
})
ET.SubElement(leaf, "title").text = "Fictional practice document"
xml_bytes = ET.tostring(root, encoding="utf-8", xml_declaration=True)
(folder / "practice.xml").write_bytes(xml_bytes)
checkpoint("baseline", inspect(xml_bytes, folder), "PRACTICE_CHECKS_PASS")
changed_reference = deepcopy(root)
changed_reference.find("leaf").set(HREF, "documents/missing.txt")
checkpoint("missing reference", inspect(ET.tostring(changed_reference), folder),
"MISSING_FILE")
payload_path.write_bytes(b"Practice document version two\n")
checkpoint("changed bytes", inspect(xml_bytes, folder), "CHECKSUM_MISMATCH")
payload_path.write_bytes(PAYLOAD)
checkpoint("truncated XML", inspect(xml_bytes[:-5], folder), "MALFORMED_XML")
checkpoint("reset", inspect(xml_bytes, folder), "PRACTICE_CHECKS_PASS")Run it with:
python3 inspect_backbone.pyThe expected output is exactly:
baseline: PRACTICE_CHECKS_PASS
missing reference: MISSING_FILE
changed bytes: CHECKSUM_MISMATCH
truncated XML: MALFORMED_XML
reset: PRACTICE_CHECKS_PASSThe five lines are the completed exercise record. The program raises an error if any observed result differs from its stated checkpoint. Keep the script and your output together if using this lesson in a team training session; there is no downloadable worksheet required to complete it.
Explain each checkpoint before applying the lesson
Baseline: several narrow checks agree
The first line means that the parser read the XML, the practice record matched this lesson's assumptions, the target file existed, and its bytes produced the recorded checksum. The title also contained text. That is useful evidence about this small fixture.
It does not establish that practice is an allowed eCTD root element, that a text file belongs in a particular submission section, or that the record is meaningful regulatory content. Those checks were never attempted. Python's ElementTree documentation explains its XML parsing and namespace handling; a successful parse is not DTD validation.
Missing reference: valid XML can describe a nonexistent file
The second checkpoint changes only the path in a copy of the XML tree. The title, operation and checksum remain unchanged. Parsing still succeeds, but there is no documents/missing.txt in the fixture directory.
This isolates a useful investigation question: did the file move, or did the reference change incorrectly? In a real controlled workflow, inspect the intended approved source before choosing a repair. Creating an empty file with the expected name would remove the missing-file symptom while leaving the actual document absent.
The practice checker resolves paths relative to its own fixture directory. A production investigation must resolve each reference from the correct XML location and account for the permitted reference forms. Do not paste a workstation's absolute path into the dossier to make a local check pass.
Changed bytes: an unchanged filename does not mean unchanged content
The third checkpoint overwrites the payload with version-two text while retaining the original XML. The file remains present and the reference still resolves. The digest comparison fails because the bytes changed.
That illustrates why checking only filenames or visible titles misses some changes. Even an apparently minor edit can alter file bytes. The correct next action depends on which version was approved: restore the intended file, or approve and rebuild the package from the intended revision. Updating a checksum without resolving that question only makes the reference agree with an unidentified change.
The script uses MD5 to illustrate this established eCTD mechanism, not to assert protection against malicious modification. Python's hashlib documentation distinguishes hash use and documents the usedforsecurity argument. Your operating environment may impose additional cryptographic restrictions; if the digest function is unavailable, record that execution limitation instead of inventing a successful result.
Truncated XML: structure must be readable first
The fourth checkpoint removes bytes from the end of the XML. The parser fails before it can inspect a leaf. Investigating a checksum at that point would skip the earlier blocking problem.
This is why a copied XML excerpt should never be presented as a runnable full document. Opening and closing tags, namespace declarations, quoted attributes and encoded text all matter to the parse. The lesson creates a complete document using an XML library so the failure is intentional and reproducible.
Reset: recovery is observable
The final checkpoint restores the original payload and reuses the unchanged original XML. Its return to PRACTICE_CHECKS_PASS shows that the introduced faults were contained and reversible.
Rerunning the script starts from a new temporary directory and repeats all five checkpoints. If your output differs, compare the saved script with the complete block before changing its expected values. Replacing the expected result with whatever appeared defeats the purpose of the exercise.
What this exercise deliberately cannot validate
| Layer | What evidence a real workflow still needs |
|---|---|
| DTD or schema | The exact applicable definition and a validator that checks conformance to it |
| Regional metadata | Correct authority, application context, controlled values and regional implementation |
| Document suitability | Approved content, actual file type, location, granularity and relevant document checks |
| Lifecycle | Existing history and correct relationship to the affected leaf or document |
| Package integrity | All applicable files, references, checksums and technical rules across the package |
| Delivery | The proper transmission route and actual acknowledgment/receipt evidence |
The lesson accepts only one new operation and one literal local path form. It deliberately rejects other operations and paths as outside practice scope. That rejection is not a statement that those forms are universally forbidden in eCTD.
For actual FDA 3.2.2 work, compare the report against the FDA validation criteria, version 4.6, rather than treating the lesson's diagnostic labels as agency error codes. The labels MISSING_FILE and CHECKSUM_MISMATCH are our teaching outputs. Record the real rule identifier, severity, applicable rule-set version and affected object from your publishing environment.
Turn a finding into a controlled correction
Use a short investigation record to connect a technical result with the publishing decision. For a fictional missing-file finding, a complete entry could read:
| Field | Illustrative entry |
|---|---|
| Observed result | Candidate package references an absent overview file |
| Evidence | Retained report, XML location, leaf ID and unresolved path |
| Intended content | Approved overview revision identified in the source repository |
| Investigation | Compare that source's approved handoff record with the candidate package |
| Decision | Rebuild the candidate from the confirmed approved source; preserve the failed candidate as investigation evidence |
| Closure evidence | New package identifier, rerun report, resolved reference and reviewer confirmation |
The table is a proposed working record, not a regulator-prescribed form. Its purpose is to prevent “fixed” from becoming an unsupported status. A report can show that the reference now resolves; the reviewer must still establish that it resolves to the intended approved content.
Also inspect neighboring references affected by the same change. A renamed directory can break several leaves even when the first report highlights only one. A replaced source can require regenerated checksums elsewhere. Retest the rebuilt candidate with the applicable complete rule set and retain the original finding alongside the closure evidence. That links the correction to the observed problem without silently rewriting the submitted history.
Interpret lifecycle instructions using the submission history
The 3.2.2 operation vocabulary includes new, replace, append and delete. A modifying operation points to an affected leaf through modified-file; replacing a filename on disk is not the same operation. Consult ICH Appendix 6 and the applicable regional rules before choosing the instruction.
Consider a fictional approved overview that has already been submitted. A new revision appears in the source repository. Before publication, the operator needs to identify the existing submitted leaf, determine the intended lifecycle action, and preserve the earlier sequence. The authoring revision number alone does not answer those questions.
The tutorial has no prior sequence and therefore cannot prove that a replacement target is current or even exists. Expanding it into a lifecycle validator would require a different fixture, historical records and much broader checks. Use the eCTD lifecycle guide for the operational distinctions rather than adding an untested replace attribute to this example.
Keep the v4.0 boundary explicit
FDA's eCTD v4.0 Technical Conformance Guide, version 1.5, June 2026 describes submissionunit.xml as the message organizing ICH and regional sequence content. A renamed 3.2.2 backbone is not that message.
When investigating a file someone calls “the eCTD XML,” first record the authority and implementation version. Then obtain the corresponding specifications, controlled vocabulary and validation configuration. Use the v4.0 transition guide, FDA eCTD resources, and ICH v4.0 resources to follow the correct branch.
For a supported FDA v4.0 workflow, book an Assyro demo with the application context and output you need to inspect. Assyro does not currently provide 3.2.2 publishing. For this lesson, the useful next step is simpler: retain the five checkpoint outputs and explain which validation layers each result leaves unresolved.
About the author
Assyro Team
Expert regulatory operations consultants helping pharmaceutical companies navigate complex compliance challenges.

