← All Journal articles

Property Operations

Stop Making Your Best Employees Babysit AI

An approval queue can preserve the very bottleneck an AI agent was meant to remove. Give routine work defined authority, evidence and escalation rules—and make human review earn its place.

Editorial illustration of a plum approval stamp and blank ticket cards beneath a gate-shaped shadow on a coral and cream background.

A crowded approval queue can feel responsible. Every action gets a human name beside it. It can also keep your strongest property managers occupied with routine work the agent was supposed to finish.

An approval queue can become a bottleneck with an audit trail.

The argument here is not that property management should remove people from every AI-assisted workflow. Copilots can be useful. Explicit approvals are necessary for consequential decisions. But blanket approval queues can preserve the delay and expert workload of the old process while delivering very little independent judgment.

A Click Does Not Prove Review

The control question isn’t whether a manager touched an action. It’s whether that manager had enough context, time, authority, and evidence to change the outcome when it mattered.

A September 27, 2026 monitoring assessment from METR breaks effective oversight into a chain: the relevant activity must be covered, the action must be visible, the monitor must identify a concern, and the human review must actually stop or redirect the right action. The assessment rated the evidence for reliable human review as weak; it did not directly test reviewers’ ability to separate real concerns from false alarms.

That is agent-safety evaluation evidence, not a property-management finding. The useful question for operators is whether the review produces an informed, independent decision. The interface alone cannot answer it.

If the only test of oversight is “a person clicked approve,” you have tested attendance, not judgment.

Repetition Can Turn Review Into Clearance Work

The concern gets sharper when the queue becomes routine.

A June 2026 preprint on AI-agent code review tracked 11,429 reviews by 400 repeat reviewers. Over time, approvals increased, elapsed time to review submission rose, and inline comments declined. The researchers said the pattern was consistent with declining scrutiny under workload, but they could not establish causation or account for possible changes in the quality of the agent-generated code. It was a software-development setting, not leasing, maintenance, accounting, or resident service.

Separately, Anthropic reported high approval rates for Claude Code permission prompts and described declining attention as prompts accumulated. That is vendor telemetry from a coding product, not a property-management benchmark.

Neither source tells you what approval rates or error patterns to expect in a portfolio. Together, they support a practical hypothesis worth testing locally: repeated, low-value prompts may drain attention without improving scrutiny.

The manager’s attention is finite. A queue full of routine work still consumes it. The rare exception that deserves expert judgment arrives in the same visual pile as every harmless reminder and status update.

Define What the Agent Can Finish

Give the system a specific job and enforceable limits.

A good boundary defines what the agent may do, which records it must verify, who it may contact, how often it may act, and which outcomes it is prohibited from creating. The agent handles an explicitly permitted action with limited consequences. A human handles exceptions and consequential decisions.

Consider a hypothetical maintenance follow-up. An agent could request a missing completion photo from an assigned vendor only when it has verified the work order, verified the recipient, checked for a duplicate request, and observed a defined reminder interval. It should carry the supporting record with the request.

That authority should stop well before the consequential edge. The agent does not approve an invoice, change a balance, authorize access, declare the repair complete, or close the work order.

This is not a claim about what any particular property-management platform supports. It is an operating design. The point is to remove routine approval clicks by constraining the action before it happens, not by asking a senior employee to bless each low-risk repetition.

Escalate the Exception With the Evidence Attached

A human escalation should arrive because something is unusual, conflicting, or consequential, not because the system generated another task.

The reviewer needs a short, legible packet: the proposed action, source records checked, the condition that failed or conflicted, what the agent did not do, and the available choices. That gives the person a real chance to evaluate and redirect the work.

I would retain explicit human judgment for money movement, balance changes, access, resident safety, habitability, legal and fair-housing implications, disputed facts, and irreversible status changes. Exact approval requirements depend on the jurisdiction, policy, and workflow. These are places where a person’s judgment can materially alter the result.

Keep Copilots Where Judgment Is the Work

Not every useful AI workflow should become an agent workflow.

A copilot can draft a vendor message, assemble maintenance history, compare documents, or surface missing information while the employee makes the decision. That is often the right model when context is messy, facts need interpretation, or the communication itself requires judgment.

In a 2021 experiment, interventions that prompted participants to think more carefully reduced overreliance compared with simple explanations. Participants rated the most effective designs less favorably. That was an experimental task, not a property-management deployment. The lesson is not to add friction everywhere. It is to add the right friction where independent evaluation matters.

A manager reviewing a disputed charge or a habitability issue needs room to investigate, question, and choose. A manager approving the fiftieth identical request for a missing photo probably needs a better-designed workflow.

Audit Outcomes, Not Just Approval Volume

Start small. Choose a narrow action with limited consequences and clear prohibitions. A sent message cannot reliably be undone; verify the recipient and contain the scope before sending. Run it under local policy, review the results, and expand only if the evidence supports expansion.

Measure whether the control changes outcomes: human review time, queue delay, approval streaks, rejection and revision patterns, rework, missed exceptions, and sampled downstream outcomes. The NIST AI Risk Management Framework calls for clearly defined human roles and says the frequency and rationale for overrides may be useful deployment information.

There is no universal approval-rate threshold or review-time target in the available evidence. The sources reviewed here do not provide comparable property-management field evidence for the proposed workflow. Its effects must be tested locally, with the actual policies and consequences at stake.

Your strongest operator should be setting standards, examining failures, and resolving exceptions. Give that person evidence and authority to intervene. Make the software earn permission to finish routine work.

A growing pile of approvals can mean you kept the delay, kept the labor, and lost the scrutiny. Make routine approvals earn their place.

Keep exploring

Ask Jane about this article.

Connect the article to GroundHaven consulting or current market evidence.

Continue with Jane
AI Management Consulting

Put perspective to work.

Discuss your priorities and budget. Build a plan around the support your team needs.

Explore consulting options