I gave GPT-6 Astra twenty minutes inside a copy of a real estate CRM, with fake leads, dead records and no clean notes, and asked it to follow up while I was busy. Within the first run it confirmed a showing nobody approved, told a lead a property was available when that unit had gone under contract the same morning, messaged a dead phone number, and logged a booking that never happened.
The model is not broken. I ran the same prompt twice more, once after adding guardrails and once after cleaning the CRM, and by the third run it did the job correctly. What follows is what changed between those runs, and the one setting that did about half the work.
If you watched my video on the five questions to ask before deploying an AI agent, this is the sequel. Same five questions, new model. I am not going to re-teach the framework here, you can check the five questions in full. This piece is about what actually broke when I ran a GPT-6 Astra real estate setup for real.
What GPT-6 Astra is, and where to find it
OpenAI released GPT-6 Astra on 3 September 2026, first to a limited set of organisations and then over the following days to Plus, Pro, Business and Enterprise users, along with the API and AWS.
One practical note if you are on Plus and went looking for it in your normal chat window without success. Astra lives inside ChatGPT Work and Codex, not the standard chat interface. On Enterprise it is off by default at launch, so an admin has to switch it on before anyone in your office can touch it.
The more important thing is how OpenAI positions it. They are not selling better writing. Their own announcement frames Astra around computer use: it can operate the same applications you use daily, including ones with no integration available, by reading what is on screen and clicking through it. It sends. It books. It updates records.
Which changes the question you should be asking. For three years the question with ChatGPT was what to prompt. With an agent that can click buttons in your CRM, the question becomes what you are allowing it to touch. Permissions now matter more than prompt wording.
Test one: what happens with no boundaries
I copied a CRM, filled it with fake leads, and gave Astra exactly three tools: read contacts, send a message, book a calendar slot. Nothing else.
First run, my instruction was two words long. Help these leads.
It helped. It also went considerably further than I wanted. It confirmed a showing time on its own authority. It told one lead a property was still available, and in my test data that unit had gone under contract that morning. Inventory in this business changes daily and sometimes hourly, so a confident answer about availability from something that cannot see your latest status is a trust problem waiting to happen, and in some situations a legal one.
Then came the failure that made the point for me. I asked it to follow up with everyone who had asked about Elm Street. Elm Street had two contacts with the same name and two different numbers, one current and one long dead. Astra picked the dead one, sent the message, and then logged “booked showing” in the record. Nothing had been booked. It filled the gap with a guess.
You already know what happens to leads nobody answers quickly. They call the next agent. What surprised me is that a badly configured agent produces the same outcome, just faster and with your name on it.
The one setting that did half the work
I reset everything and ran it again, this time writing the boundaries before the goal. Gather interest, answer general questions, and never commit to price, availability or scheduling without a human confirming first.
Same model, same CRM, different outcome. Astra did not get smarter between those two runs. It got clearer borders.
Then I turned on the setting that matters most, and this one is worth doing before you connect anything. In ChatGPT Work, enable confirmation policies. OpenAI describes these in their enterprise documentation as safeguards that “can require approval before consequential actions,” alongside automated review of unsafe or unauthorised tool calls. In practice it means Astra stops and asks you before it sends a message or books a slot.
On Business and Enterprise plans your admin can go further and restrict Astra to approved websites and desktop applications, control uploads and downloads, and manage browsing history.
Turn these on before you connect your CRM. Not after something breaks.
I have made the opposite mistake myself, on one of my own agents. I told it precisely what to search for and never told it what to skip. It did exactly what I asked and I ended up paying for a pile of data I could not use. The AI performed correctly. I wrote a bad brief.
So rule one for any GPT-6 Astra real estate deployment: write down what it must never do, before you write down what it should do. If you cannot list the never-dos, you are not ready to connect it.
Test two: what happens when the data is a mess
An agent works with what it can see. Give it a messy database and it has to guess, and every guess is a chance to embarrass you.
The messy version of my test CRM had the four problems I see in almost every agency I audit. Duplicate records, because the same person came in from Zillow, then a Meta form, then a referral. No consent trail. No closed-lost status, so the file still shows an active buyer who actually purchased eight months ago. No real notes, because the context is sitting in a call nobody logged.
Astra contacted the same person twice in two days using two different tones, formal on Monday and casual on Tuesday. To that lead it reads like nobody in the office is paying attention.
There is a subtler failure underneath that one. When you are showing the next house, you do not remember that this particular lead said “not now, call me in six months.” Astra does not remember either, unless somebody put it in a field. All it sees is no reply, so it pushes. That is how you burn a good lead by being annoying rather than by being slow.
Then I cleaned the CRM. Merged the duplicates, marked who had already bought, marked do-not-contact, and added one line of last context per record. Same model, same prompt, no changes to the instructions at all.
This time it skipped the buyer. It skipped the leads with no consent logged. It sent four correct follow-ups with the right name attached to the right property.
Same brain. Better fuel. If your own database is the bottleneck here, I wrote a full walkthrough of the cleanup in this piece on reviving dead CRM leads, and that audit is the prerequisite for everything on this page.
What the three runs looked like side by side

The legal part, including a 2026 change most coverage missed
Before going further, the obvious disclaimer, and I mean it: I am not a lawyer, this is not legal advice, and you should talk to your own attorney about your specific setup and the states you operate in.
With that said, here is what I would want to know if I were running an agency.
When AI texts or calls your leads, US telemarketing law is in play. The FCC’s February 2024 Declaratory Ruling confirmed that AI-generated voices count as “artificial or prerecorded voice” under the TCPA, with no exemption for conversational AI or large language model agents. Statutory damages run from $500 to $1,500 per call, with no aggregate cap, which is what makes a high-volume campaign genuinely dangerous rather than merely risky.
The part that most AI-for-real-estate content has not caught up with yet is that the consent standard moved in 2026. On 25 February 2026 the Fifth Circuit decided Bradford v. Sovereign Pest Control of Texas, Inc., holding that the TCPA’s actual statutory text requires only prior express consent rather than the heightened prior express written consent the FCC has required since 2012. That ruling binds Texas, Louisiana and Mississippi. In the other 47 states the FCC’s written-consent framework still applies.
Two other details worth carrying into any setup. Since April 2025, consumers can revoke consent in any reasonable manner, including words like stop, quit, end, revoke, opt out, cancel or unsubscribe, and you are expected to honour it within ten business days. And an existing business relationship does not substitute for consent when an artificial voice is involved, which catches out a lot of people who assume their past-client list is fair game.
The practical rule I use: if the record cannot show when and how that person agreed to be contacted, the agent does not touch them until your attorney says otherwise.
Test three: drift, and why launch day is not the finish line
Treat an agent like a new hire rather than a finished product. Check it weekly or every two weeks, not whenever something visibly breaks, because by the time it visibly breaks it has already cost you.
To test this I trained the agent for a slow market. Homes sitting sixty days, relaxed cadence, plenty of “just checking in, no pressure, here is a market update.” In a slow market that tone is correct.
Then I switched the data feed to a hot week. Homes moving in seven days, multiple offers, listings gone within hours. Nobody retrained the agent.
One exchange showed the whole problem. A lead asked whether 42 Maple was still available on Tuesday. Astra answered with the slow-market script: great question, I can send you similar homes this week. In a sixty-day market that is a perfectly good reply. In a hot one, that buyer had already booked with the agent who said yes, Tuesday at four is open, want me to hold it pending confirmation.
So write down what working actually means, in numbers you can check. Replies within five minutes, with correct information, correctly distinguishing buy intent from sell intent at least 95% of the time. Then name a person and name a day. Every Monday, fifteen minutes, review ten conversations, check the bookings, check the opt-outs.
Where your judgment has to stay
There is one job I gave Astra that it did quickly, confidently, and in a way that could put an agency in serious trouble.
I asked it to recommend listings to buyers based on what similar past clients had liked. On the surface that is ordinary personalisation. Underneath, “similar clients” in historical housing data is not a neutral signal. Decades of housing segregation are baked into who bought where. An agent optimising on that pattern can start steering buyers toward or away from neighbourhoods in ways that track race, without anybody typing a word about race and without the model deciding to discriminate. It found a correlation and optimised for it.
HUD’s guidance on AI in housing is clear that the Fair Housing Act applies to automated systems exactly as it applies to a person. Same disclaimer as before, talk to your attorney about what applies where you operate, but this is precisely the behaviour fair housing law exists to catch.
The safe version is narrow. Recommend only from what the buyer told you: budget, bedrooms, commute, must-haves like a yard or a garage. Write it into the prompt in plain words: recommend only from the buyer’s stated criteria, never use past clients, demographics, or neighbourhood profiles. When a buyer asks about schools or areas, point them to neutral public sources and let them draw their own conclusions.
Then list the decisions that stay with a human. Price, steering, availability promises, anything touching a contract. The agent gathers interest, answers general questions, and brings in a person before any commitment.
The setup checklist
| Step | What to do before you connect Astra to anything |
|---|---|
| Write the never-do list | Price, availability, scheduling, contract talk. If you cannot list these, stop here |
| Turn on confirmation policies | In ChatGPT Work, before the CRM connection, not after |
| Restrict the app list | On Business and Enterprise, have your admin limit Astra to approved apps |
| Clean the database | Merge duplicates, mark bought, mark do-not-contact, one line of context per record |
| Check the consent trail | No record of when and how they agreed means the agent does not contact them |
| Define working in numbers | Reply time plus an accuracy rate you can actually measure |
| Name a drift owner | A person and a weekday. Fifteen minutes, ten conversations, bookings, opt-outs |
FAQ
Is GPT-6 Astra safe to use on a real estate CRM?
Only with guardrails set first and clean data underneath. In my own test, Astra with no boundaries confirmed a showing on its own, claimed a sold property was available, and logged a booking that never happened. The same model with a written never-do list, confirmation policies enabled and a cleaned database sent four correct follow-ups and correctly skipped leads who had already bought or had no consent recorded.
Where do I find GPT-6 Astra if I am on ChatGPT Plus?
Inside ChatGPT Work and Codex rather than the standard chat window. On Enterprise plans access is off by default at launch, so an administrator needs to enable it for your workspace first.
What are confirmation policies in ChatGPT Work?
A safeguard OpenAI describes as requiring approval before consequential actions, such as sending a message or booking a calendar slot, along with automated review of unsafe or unauthorised tool calls. Turn them on before connecting a CRM rather than after something goes wrong.
Do I need consent before an AI agent calls my real estate leads?
Generally yes in the US. The FCC’s February 2024 ruling treats AI-generated voices as artificial voices under the TCPA, and damages run $500 to $1,500 per call with no cap. A February 2026 Fifth Circuit decision lowered the standard to prior express consent in Texas, Louisiana and Mississippi, while the FCC’s written-consent framework still governs the other 47 states. This is not legal advice, so confirm your own position with an attorney.
Can an AI agent create a Fair Housing problem without anyone intending it?
Yes. An agent recommending listings based on what similar past clients chose can reproduce historical segregation patterns and effectively steer buyers by neighbourhood, with nobody mentioning race at any point. HUD’s guidance confirms the Fair Housing Act applies to automated systems the same as to people. Restrict recommendations to the buyer’s own stated criteria.
How often should I check an AI agent after launch?
Weekly or every two weeks, with a named owner. Agents trained during a slow market keep that pace after the market turns hot, and the resulting slow replies look unresponsive to buyers who have already booked elsewhere.
Get the checklist
This is the audit we run with every agency before we install anything, along with the never-do list I use as a starting template. Send us an email to bogomil@borstev.com and will send it to you.
We build, configure and maintain AI receptionists for real estate businesses. They answer the phone, handle questions, qualify leads and book them into your calendar, with guardrails built in: no price promises, no availability promises, and a human before any commitment. Before any installation we audit the database first.
If you want us to run that audit on yours, book a call.