Skip to content

Source Document Import

Source Document Import frees the coordinator from re-typing source documents into CRFs. You upload a subject’s document batch — scans and photos of paper forms, PDFs, DOCX summaries, XLSX/CSV lab tables, plain text, mixed together — and the system routes the pages to the visits and forms of the casebook, extracts the values and prepares a proposal covering several CRFs at once.

0:000:00
From a visit diary scan to a saved value: analysis, a conflict with the entered value, a reason for change, apply
  1. Batch. In the “Source import” workspace you create a batch for a specific subject and add documents — any number, any supported format.

  2. Analysis. The system reads the documents (scans go through recognition), routes pages to the visits and forms of the schedule, extracts values and checks them against the data already entered. Progress is shown stage by stage.

  3. Review. Every value comes with its recognition confidence, the exact source quote and a comparison verdict: new, matches (already in the form — nothing to overwrite) or conflict (the form holds a different value). Clicking the quote opens the document page right next to the data.

  4. Decisions. A value can be accepted, corrected or rejected. An SDV-critical field cannot be accepted without opening the source. Accepting over a conflict requires a reason for change — like any edit of a saved value in an EDC.

  5. Apply. One gesture writes the accepted values into the CRFs — under your name. Missing visits and forms are created along the way. The report honestly lists everything the core refused (a locked form, for example).

Every proposed value keeps a complete trail: which document it came from (with a checksum), which page and which quote, at what confidence, who accepted it, when and with what reason, and what was finally written. The audit record carries the provenance marker “proposed by a module — confirmed by a human”. The expandable “How this result was produced” panel shows the model and analysis timings.

  • Never writes data itself — only a proposal; applying is always the human’s act.
  • Never sees blinded fields — they are excluded before analysis.
  • Never invents: a value without a clear source stays empty, an unreadable one is flagged, and a disagreement between two documents requires an explicit choice.
  • Checks identity: if a different subject’s code appears in the documents, the batch gets a warning.