The Aperture Field - The Human System in the Loop Graphic - understanding the relationship between HITL and the system within which it operates

Hi, I’m Monet Goode.

From Human Checkpoints to Organizational Learning in Responsible AI

The Human System
in the Loop

From Human Checkpoints to Organizational Learning in Responsible AI

Kari LaMotte | The Aperture Field
Working Paper · October 2026

ABSTRACT

Human-in-the-Loop (HITL) approaches to Responsible AI often position human oversight as a safeguard between an AI system and a consequential decision: an AI produces an output, a human reviews it, and a decision proceeds. This model is useful, but incomplete. It can imply that the presence of a human reviewer is itself evidence of meaningful oversight and that failures originate primarily in the AI system or its application.

This paper proposes a broader unit of analysis: the human system in the loop.

AI operates within existing sociotechnical systems shaped by organizational rules, policies, historical practices, incentives, data, norms and definitions of success. A human reviewer and an AI system may therefore independently perform their assigned functions correctly while jointly producing a problematic outcome because both are operating within the same flawed governing frame.

Responsible oversight must consequently do more than catch AI errors. It must create pathways through which consequences, monitoring and challenges from affected people can travel upstream to roles with sufficient authority to examine the assumptions and rules governing the system itself.

The distinction is between a system that repeats and a system capable of learning.

Watch this short explainer video on the Human System in the Loop framework by The Aperture Field.

THE HUMAN CHECKPOINT IS TOO SMALL

A common representation of human oversight is deceptively simple:

AI → HUMAN → DECISION → AFFECTED PERSON

The intuition is sensible. Where AI contributes to consequential decisions, a human can provide judgment, context and accountability before the decision reaches the person affected by it. But the presence of a human does not necessarily establish meaningful oversight.

At scale, requiring a person to approve every AI output can become ceremonial rather than substantive. [1] More importantly, even an attentive and capable reviewer may be unable to identify a deeper class of failure: the AI and the human may both be operating correctly according to a governing rule that is itself inadequate.

Consider a simplified, illustrative loan decision:

An organization has established criteria under which an applicant is classified as high risk. An AI system applies those criteria and recommends declining an application. A human reviewer carefully checks the application against the same organizational criteria and confirms that the AI has applied them correctly.

The decision is declined.

At the case level, neither the AI nor the human reviewer necessarily failed.

Yet the resulting decision may still reveal a problem.The question is no longer simply:

Did the AI make the decision correctly?

It becomes:

Is the rule against which both the AI and human were evaluating the case itself justified?

Human review cannot provide independent validation of a governing assumption when the reviewer is operating inside that same assumption.

This suggests that the “human” in Human-in-the-Loop has been drawn too narrowly.

From the Human Reviewer to the Human System

The individual reviewer sits within a larger human system. That system includes organizational policies, historical practices, targets, norms, authority structures, data, incentives and formal or informal rules. These conditions shape what the organization considers relevant, measurable, acceptable and successful. AI does not enter a neutral environment and begin making decisions from scratch. It enters a system that is already making distinctions.

The resulting architecture looks less like:

AI → HUMAN → DECISION

and more like a sociotechnical system in which governing rules shape both AI and human judgment. [2] This distinction matters for accountability. The unit of analysis may be the sociotechnical system, but a “system” cannot become an accountability sink. Accountability must remain attached to identifiable people, roles and organizations with actual authority, knowledge and control.

Different human functions may therefore occupy different positions in the system.

  • A case reviewer may have authority to determine whether a particular decision follows established rules.

  • A review or governance authority may have the additional authority to examine whether those rules remain appropriate.

Collapsing both functions into a single box labelled HUMAN obscures an important difference between operating within a system and examining the system itself.

Contestability Completes the Case Loop

The affected person should not appear merely as the endpoint of the decision chain. Suppose the declined applicant believes the decision failed to account for relevant circumstances. They request a review. That contest provides information back to the organization. A meaningful review mechanism must have somewhere for that information to go: to a person or function with enough information and authority to examine the case and determine an appropriate response.

Importantly, review does not imply reversal. A review may reasonably produce at least three outcomes:

  1. Uphold. The original rule and its application withstand scrutiny, and the original decision remains.

  2. Exception. The rule remains appropriate generally, but the particular case warrants different treatment.

  3. System review or change. The case, potentially alongside monitoring or other evidence, exposes a reason to question the governing rule, threshold, category or assumption itself.

In each case, the accountability loop should return to the affected person with the outcome and the reasoning behind it. This creates a complete case-level relationship:

Decision → Affected Person → Contest → Review → Outcome + Reasoning → Affected Person

Contestability is therefore not simply an intake mechanism through which an organization extracts useful information from affected people. It is reciprocal: the organization receives a challenge, exercises judgment, and provides an intelligible response. [3]

Nor does every contest need to trigger system change. The important design property is that evidence can travel far enough to reach the level at which the relevant question can actually be considered.

A System That Repeats

AI systems are frequently introduced into processes with substantial histories. Past organizational decisions may become data. Existing rules may become thresholds, objectives, categories or other features of an automated process. New decisions are made, recorded, and may subsequently become part of the informational environment shaping future decisions.

This can produce a repeating system. The risk is not limited to an AI inventing a new bias. AI can also scale, stabilize or obscure existing patterns. [4]

A human reviewer can remain present throughout this process. If that reviewer is primarily verifying whether AI decisions conform to the same governing rules, however, the human checkpoint may do little to interrogate the assumptions generating those decisions. The system can therefore contain human oversight while continuing to reproduce itself.

This raises a different diagnostic question from asking whether a problem was “caused by AI”:

If we remove the AI, what changes in the magnitude, speed, visibility and contestability of the problem?

AI may create some harms, amplify others, or make pre-existing problems newly visible. Determining causal origin matters, but it does not by itself resolve accountability for what happens when an organization automates and scales a process.

A System Capable of Learning

A learning system contains additional feedback pathways. Monitoring may reveal a pattern across many decisions. An affected person may surface information invisible to aggregate data. Case reviews may expose recurring exceptions. These signals may indicate a problem with the AI model itself. In those circumstances, correcting the AI is appropriate.

But feedback must also be capable of travelling farther upstream. Where evidence suggests that the AI is faithfully executing a problematic rule, someone with appropriate authority must be able to ask:

Is the rule itself still justified?

In some cases, inquiry may need to travel farther upstream still. A problematic outcome may not originate in a particular rule or model, but in the way the organization has systematized the activity itself. Responsible oversight therefore requires the capacity to question not only how an automated process operates, but whether the activity should be automated or structured as a standardized decision process at all. The appropriate response to evidence may sometimes be to redesign the process, remove automation from it, or reconsider the underlying systemization itself.

This creates two forms of feedback rather than substituting one for another.

  • One pathway corrects performance within the existing governing frame.

  • Another allows the organization to examine the governing frame itself.

This distinction has an established lineage in organizational learning, particularly the distinction between single-loop and double-loop learning associated with Chris Argyris and Donald Schön. [5] The contribution proposed here is not the invention of that distinction, but its application to Human-in-the-Loop governance.

Much HITL thinking asks whether a human can identify and correct failures in AI-mediated decisions. The Human System in the Loop asks an additional question:

Can evidence from AI-mediated decisions cause the human system to examine and, where warranted, revise the rules by which both humans and AI are operating?

A system capable of learning does not guarantee that the rule changes. It guarantees that consequential information has somewhere meaningful to go.

Oversight at the Appropriate Level

This broader model also complicates the assumption that responsible oversight means human review of every output.

Human judgment is a limited and valuable system resource. Requiring repetitive confirmation of high-volume outputs may consume that capacity without adding meaningful information or discretion. Different risks require different forms of oversight.

Case-level review may be appropriate where a reviewer possesses context the system lacks, where individual errors carry substantial or irreversible consequences, or where human answerability is itself important. Other functions may be better served through sampling, exception handling, monitoring, audit, contestability or periodic review of governing rules.

The question is therefore not:

Where can we insert a human?

but:

Where does human judgment add value, what authority does it require, and how does information move between those levels of judgment?

This also introduces a temporal concern. Human capacity is not necessarily static. Repeated reliance on AI may change the knowledge, attention or discriminatory ability on which oversight depends. [1] A governance architecture that was meaningful at deployment may become nominal later if responsibility remains assigned to a role whose actual capacity to detect and correct failure has changed.

This dynamic oversight gap is a secondary implication of the model and warrants further investigation.

The Human System in the Loop

The central shift proposed here is therefore relatively small in language but substantial in architecture. Human-in-the-Loop asks us to keep a person meaningfully involved in AI-mediated decisions.

The Human System in the Loop asks us to examine the larger system that makes that person's judgment possible and gives it meaning. A responsible system should be capable of:

  • applying and evaluating decisions at the case level;

  • receiving challenges from affected people;

  • returning decisions and reasoning to those people;

  • detecting patterns across decisions;

  • correcting AI failures when they occur;

  • distinguishing case exceptions from systemic problems;

  • routing evidence to identifiable roles with appropriate authority; and

  • examining and revising governing assumptions when evidence warrants it.

The distinction is ultimately not between human decision-making and AI decision-making. It is between a system that can only repeat and a system capable of learning. Keeping a human in the loop is not enough if the loop itself cannot question the system that created it.

Responsible AI requires a human system capable of learning from the consequences it produces.

Limitations and Open Questions

This framework is intentionally conceptual. It does not specify universal thresholds for when individual human review, automated monitoring, sampling or other forms of oversight should be used. Those choices depend on the severity and reversibility of potential harm, the detectability of failure, the validity of available ground truth, the distribution of impacts, and the information and authority available to human reviewers.

Several questions remain open.

  • How should organizations determine when a pattern of individual contests warrants systemic review? 

  • How can affected people meaningfully contest decisions without creating inaccessible or burdensome appeal processes? 

  • What forms of explanation constitute adequate reasoning? 

  • How should organizations monitor whether oversight capacity itself is changing through repeated AI reliance? 

  • And how should accountability be distributed across developers, deployers, managers, reviewers and governance authorities without allowing distributed responsibility to become diluted responsibility?

The framework does not assume that learning necessarily produces improvement. Feedback can reinforce existing assumptions as easily as challenge them. The purpose of the learning architecture is therefore not to guarantee correct outcomes, but to preserve the capacity for consequential evidence to reach the places where assumptions can actually be examined.

References

[1] Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779.

Foundational analysis of monitoring, skill maintenance and the tasks left to human operators. This supports the oversight concern; the AI-specific dynamic oversight gap remains a proposed implication.

[2] National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1.

See the Executive Summary; GOVERN 2 and 3; MANAGE 2 and 4; and Appendix C. These address sociotechnical context, accountable roles, human–AI interaction, alternatives and continuing risk management.

[3] National Institute of Standards and Technology. (n.d.). NIST AI RMF Playbook.

See GOVERN 3.2, MEASURE 2.8 and 3.3, and MANAGE 2.1 and 4.1 for oversight, feedback, appeals, non-AI alternatives and post-deployment monitoring.

[4] Schwartz, R., Vassilev, A., Greene, K., Perine, L., Burt, A., & Hall, P. (2022). Towards a Standard for Identifying and Managing Bias in Artificial Intelligence. NIST SP 1270.

See Section 2.1.2 on systemic bias and Section 3 on sociotechnical sources of bias.

[5] Argyris, C., & Schön, D. A. (1978). Organizational Learning: A Theory of Action Perspective. Addison-Wesley.

Foundational treatment of single-loop and double-loop organizational learning.