← All Journal articles

Property Operations

AI Can Build the Tool. Property Managers Still Have to Specify the Work.

AI may lower the barrier to building internal property-management tools, but it does not define the work for you. Re-map leasing, maintenance, collections, inspections, and reporting with clear permissions, exceptions, acceptance tests, and human approval points.

An ochre pencil draws a line across a navy architectural plan beside a doorway rising from the page.

The Prototype Barrier Is Lower. The Operating Bar Is Not.

The build-versus-buy question is changing for property-management operators.

On September 3, 2026, OpenAI said its GPT-6 Astra model could perform multistep professional work, use computers and browsers, and create hosted web applications from prompts. In a June 2, 2026 report, OpenAI also said non-developers made up about 20% of Codex users and that its own nontechnical teams were using Codex to create internal apps and dashboards.

Those are vendor-reported claims, not proof that a property-management workflow is ready for production. Taken together, they suggest that a dedicated technical team is no longer the only route to an internal prototype. They do not establish how much technical review, integration work, or ongoing support a particular prototype will require.

For property-management operators, the likely constraint shifts toward explaining how the work actually gets done: which record controls when systems disagree, what facts must exist before a decision, which exceptions require escalation, who can approve an action, and what evidence must remain in the file.

Operators should not assume a model will correctly reconstruct undocumented rules. A polished workflow can still mishandle an incomplete lease file, a possible duplicate work order, an unusual charge, or a resident communication requiring judgment.

This is a reason to re-map core workflows. It is not a reason to hand them over wholesale.

What Easier Tool Creation Changes

The near-term opportunity is not necessarily an autonomous leasing, maintenance, or collections system. A more controlled starting point is a narrow tool that reduces administrative work while a person retains responsibility for the final decision or external action.

Consider work with four characteristics:

  • A clear trigger, such as a new inquiry, completed inspection, incoming work order, or month-end reporting cycle.
  • A bounded output, such as a draft message, categorized request, extracted document field, exception queue, or variance summary.
  • A defined source of truth.
  • A reviewer who can recognize a bad result before it changes a record or affects a resident, owner, or vendor.

That combination can make a controlled pilot plausible even without in-house developers. This is an operating hypothesis, not a property-management-specific research result.

The reverse is also true. A workflow should not gain authority simply because a model can navigate software or generate an application. Lease changes, charge postings, collection notices, vendor authorizations, and other external commitments call for documented decision rights, constrained access, testing, approval controls, and a way to investigate what happened.

Convert Operating Knowledge Into a Specification

Before an AI tool touches production data, write a workflow specification that a new team member could follow and a reviewer could test.

It does not need to resemble software documentation. It does need to answer the operating questions precisely.

Define the job and its boundary

State one job to be done.

“Help the maintenance coordinator identify work orders that may be duplicates” is a usable starting point. “Manage maintenance” is not.

For that narrow job, define:

  • The trigger that starts the work.
  • The expected output.
  • The person accountable for the workflow.
  • The point at which the tool stops and a person takes over.

Establish a source-of-truth hierarchy

Where a workflow can encounter conflicting records, the tool needs an explicit hierarchy. A unit status may appear in one system, a note in another, and supporting information in an attachment or verbal update.

The tool may be allowed to summarize approved sources but required to flag discrepancies rather than resolve them. Do not assume it will infer which system controls in every circumstance.

Write the decision table

A decision table turns experience into repeatable rules.

For a maintenance-triage pilot, specify which fields are required to classify a request, which conditions trigger emergency escalation, which conditions create a possible-duplicate flag, and which missing information requires a follow-up queue rather than a category assignment.

For leasing, identify the required facts for a draft response, the approved information sources, and the situations that must be routed to a human before any response is sent.

Name prohibited actions and escalation rules

The most useful instruction is often what the tool may not do.

Examples include:

  • Do not send an external message without the assigned reviewer’s approval.
  • Do not alter a lease record.
  • Do not approve vendor work.
  • Do not post or recommend a charge when required documentation is missing.
  • Do not resolve conflicting data; create an exception for review.
  • Do not make a decision requiring legal, fair-housing, safety, or policy interpretation without the designated human owner.

These boundaries distinguish a useful assistant from an unauthorized decision-maker.

Re-Test Each Operating Segment

Every portfolio has different systems, policies, staffing patterns, and local requirements. The examples below are operating hypotheses for controlled pilots, not evidence that AI has proven accurate or secure in leasing, maintenance, collections, inspections, or reporting. The reviewed sources contain no property-management-specific results for those functions.

Leasing: Draft and organize before communicating

A lower-authority pilot could assemble approved property information into a draft response for review, extract fields from an inquiry, or identify files missing required internal documentation.

The acceptance test is not whether the draft sounds friendly. It is whether it uses approved source material, avoids unsupported statements, identifies missing information, and routes exceptions to a reviewer.

Do not casually delegate outbound communications, lease-file changes, eligibility judgments, or policy-sensitive decisions. Those actions can create commitments or require interpretation beyond document retrieval.

Maintenance: Triage the queue, not the repair decision

Possible bounded internal tasks include summarizing work-order histories, identifying potential duplicates, preparing a coordinator’s queue, or drafting a status update for human approval.

The specification should identify emergency indicators, the data needed to distinguish a duplicate from a repeat issue, records the tool may consult, and the escalation path for incomplete or conflicting information.

A tool should not independently authorize work, direct a vendor, or close a work order merely because it identifies a likely pattern. The coordinator retains the operating decision.

Collections: Organize evidence before making contact

A tool may help prepare internal file summaries, identify missing documentation, or flag accounts for human review based on operator-defined criteria.

Treat these uses as controlled-pilot ideas rather than established practices. Messages, charges, payment arrangements, and account actions can have significant consequences. Any use in this segment needs clear approval points, complete records, and review under the requirements governing the portfolio.

Inspections: Turn observations into a reviewable queue

A pilot may extract observations from completed inspection materials, organize them by unit or category, or prepare a follow-up list for a team member.

The tool should retain the source material used for each item and flag unclear records. It should not turn an ambiguous observation into a final charge, repair authorization, or resident-facing conclusion without human review.

Reporting: Explain variances, then verify the explanation

Reporting is attractive because the output may remain internal. It still presents a risk: a confident explanation can be wrong.

Artificial Analysis reported on August 10, 2026, that the leading model in its quantitative-analysis benchmark completed 54% of tasks correctly across all five runs. Among the failures the organization classified, 57% involved anchoring on an incorrect early interpretation.

This was not a property-management benchmark, so those rates should not be applied directly to portfolio reporting. The narrower warning is relevant: a model can begin with the wrong interpretation of a spreadsheet or document and build a plausible analysis on top of it.

Use reporting tools to prepare a variance narrative, identify records requiring investigation, or compare defined fields. Require a reviewer to verify calculations, source data, assumptions, and conclusions before an analysis informs an owner, lender, or operating decision.

Use an Authority Ladder

The following authority ladder is an operator-authored governance framework, not an externally validated standard. Its purpose is to match authority to evidence and risk. Begin at the lowest useful level and raise authority only after the workflow performs reliably under repeated testing.

Level 1: Assist and draft

The tool summarizes, extracts, classifies, or drafts. A person reviews the complete output before using it.

This is often a practical first-pilot structure because the tool cannot alter records or communicate externally on its own.

Level 2: Triage and recommend

The tool prioritizes a queue or recommends a next step under documented criteria. A person accepts, rejects, or modifies the recommendation.

Log recommendations and reviewer decisions. That record helps the team assess whether the tool is useful and where its instructions need revision.

Level 3: Prepare an action for approval

The tool fills out a proposed system action or communication, but a designated person must approve it before execution.

This level calls for tight controls over accessible data and the actions the tool may prepare. The approver should see the evidence behind the proposed action, not only the model’s conclusion.

Level 4: Act under narrowly defined authority

The tool takes an action within a constrained, preapproved rule set.

This is the highest-risk level in this framework and should be exceptional, not the assumed destination for every pilot. It calls for clear permissions, logging, monitoring, defined exception handling, a shutdown path, and re-testing when policy, software, or workflow conditions change.

A Prompt Is Not a Control System

Better instructions matter. They do not replace testing and controls.

A preliminary August 2026 arXiv study examined 12 web applications generated in one non-iterative round. Researchers found 75 confirmed security findings. Explicit security requirements reduced confirmed findings from 51 to 24, but manual testing still identified the most severe issue.

The study used a small corpus and generated each variant once. Its authors described the results as preliminary and descriptive rather than statistically established. The study does not supply a universal defect rate for AI-built applications. It does support testing beyond an instruction to “be secure.”

The need for controls becomes more material when a tool can connect to operating systems or take actions. In its September 3, 2026 safety overview, OpenAI classified GPT-6 Astra at its Critical cybersecurity capability threshold under its Preparedness Framework. OpenAI said the classification concerned what the model could do with suitable tools and access.

That classification does not mean every internal property tool will create a cybersecurity incident. The relevant operating inference is that connected-system access, permissions, and deployment boundaries should not remain undefined.

Use least-privilege access as a control principle. Give a pilot only the data and actions required for its specific task. Avoid broad credentials, unrestricted file access, or system-change authority granted for convenience.

Minimum Requirements Before a Production-Data Pilot

The following checklist is an operator-authored starting point informed by the reviewed security and reliability evidence. It is not a complete technical security standard.

Before moving beyond synthetic examples or appropriately handled historical cases, require a written record of:

  • Process owner: One person accountable for the workflow and its revisions.
  • Defined job: A narrow task with a clear trigger and bounded output.
  • Approved data: The specific systems, files, and fields the tool may use.
  • Source hierarchy: Which record controls when data conflicts.
  • Permission boundary: The minimum access needed to complete the task.
  • Approved and prohibited actions: What the tool may prepare, recommend, or do, and what it must never do.
  • Exception path: Conditions requiring human review, including missing information and conflicting records.
  • Evidence retention: The inputs, output, decision, and reviewer action retained for review.
  • Human checkpoint: The named role that must approve an external communication, system change, or other consequential action.
  • Acceptance tests: Ordinary cases, edge cases, known failure patterns, and expected escalations.
  • Rollback and shutdown process: How to stop the workflow and correct affected work if a control fails.
  • Review cadence: When the team will examine performance, process changes, and needed instruction updates.

The purpose is to determine whether a tool reduces work or merely moves hidden review work downstream.

Test Exceptions Before You Trust the Happy Path

Do not test only complete and consistent inputs. Build a test set from historical cases handled appropriately for the purpose, including:

  • Missing documents or incomplete required fields.
  • Conflicting unit, resident, or ledger information across approved sources.
  • Duplicate or repeated work orders.
  • A resident message that should be escalated rather than answered from a standard template.
  • A vendor bid outside an established approval threshold.
  • A charge or inspection record with missing supporting evidence.
  • Late, corrected, or unusually formatted reporting data.

Define success before the test begins. A successful tool does not merely produce fluent prose. It uses the correct approved sources, reaches the expected output when the evidence supports one, preserves relevant evidence, and escalates when the record does not support a decision.

Run cases more than once when the workflow depends on model interpretation. The Artificial Analysis benchmark is not a property-management evaluation, but its five-run methodology and reported reliability gaps support checking whether results remain consistent. An answer that is correct once but inconsistent across similar runs may not reduce review work enough to justify production use.

Decide Whether to Build, Buy, or Document First

The lower prototype barrier does not mean custom building is always the right answer.

A custom internal prototype may fit a narrow, stable process reflecting how a specific portfolio operates. A vendor product may be more appropriate for a common or integration-heavy workflow when the organization prefers external product ownership to ongoing internal maintenance. These are decision considerations, not universal rules. Either choice can fail when the underlying process remains poorly understood.

Ask five questions before deciding:

  1. Is the workflow stable enough to describe in a decision table?
  2. Does the process involve sensitive data, high-consequence actions, or difficult approval requirements?
  3. Can the organization support ongoing ownership of instructions, testing, access reviews, and exceptions?
  4. Does the workflow need deep integration with systems of record?
  5. Will the time saved survive the reviewer workload needed to catch errors and manage edge cases?

If the workflow is not stable enough to describe, document the work before building or buying. A generated interface does not resolve an unclear operating process.

Start With One Bottleneck and One Authority Level

OpenAI’s September 3, 2026 release is a timely reason to revisit which internal tools may now be feasible. Its app-building and multistep-work statements are vendor evidence of a changing capability environment, not a guarantee about a particular property-management workflow or a permanent market leader.

As reviewed on September 9, 2026, frontier-model capabilities, access, pricing, and safeguards remain time-sensitive. A durable next move is smaller: choose one administrative bottleneck, map the actual workflow, identify decision rights and exceptions, write acceptance tests, and pilot at the lowest authority level that can create value.

A GroundHaven Portfolio Performance Review can serve as one structured starting point for mapping workflows, decision rights, exceptions, data access, and handoffs before an organization invests in an internal tool or vendor platform. This is owner-supplied service positioning, not a claim that GroundHaven builds software, provides technical security services, or guarantees a tool’s performance.

The aim is not to automate judgment out of property operations. It is to make repeatable work explicit enough that a tool can help without quietly taking on authority it was never given.

Keep exploring

Ask Jane about this article.

Connect the article to GroundHaven consulting or current market evidence.

Continue with Jane
AI Management Consulting

Put perspective to work.

Discuss your priorities and budget. Build a plan around the support your team needs.

Explore consulting options