EvidenceSheet

MS-4.1 Measurement approaches for identifying AI risks are connected to deployment contexts and informed through consultation with domain experts and other end users, and approaches are documented

Measurement approaches for identifying AI risks are connected to deployment context(s) and informed through consultation with domain experts and other end users. Approaches are documented. The measurement design is infor

4
artefacts
0
held by a system
2
at each review
hard
to go live
Policy repository / GRC workspace
where the evidence lives
teal = a system already holds it · olive = produced at each review

system holds itEvidence a system already holds

none for this control

periodic reviewEvidence produced at each review

  • Records of consultation with domain experts and end users on measurement design · Document repository
  • Evidence that consultation changed the measurement approach · Document repository

governing documentDocuments that govern the control

  • Documentation of the measurement approaches and their connection to deployment contexts · Policy repository / GRC workspace
  • Identification of the deployment contexts the measurement is meant to cover · Policy repository / GRC workspace

First move

This control is evidenced by people and documents, not systems. Put the document under version control with an owner and review date, and log each review as a record with reviewer and date. Do not try to automate it.

Common gaps auditors find

Do this for your whole sheet

Paste the rows you run your controls from and get this mapping for every control at once, with the periodic-review ones flagged and a first move per row. No account for the first run.

Build my evidence sheet

MS-3.3 Feedback processes for end users and impacted communities to report problems and appeal system outcomes are established and integrated into AI system evaluation metrics · MS-4.2 Measurement results regarding AI system trustworthiness in deployment contexts and across the AI lifecycle are informed by input from domain experts and other relevant AI actors to validate whether the system is performing consistently as intended, and results are documented