Two answers to the same Dubai property question share about half their firms
I sent 10 property questions through ChatGPT's search mode 5 times each, via a third-party scraper API, extracted the organization names from all 50 answers, resolved them against the Dubai Land Department broker register, then audited the extraction against runs read by hand. The median pair of answers to the same question overlapped at a Jaccard of 0.32. For lists of similar length, that means roughly half the firms in either answer turn up in the other.
Ask ChatGPT which agency to use in Dubai, then ask again. You get a different list. I wanted a number for how different, so I sent the same strings through a third-party scraper API 50 times and stored every response. The endpoint reports no model identifier, so this measures what that interface returned on one day, not the behaviour of any named model.
The sharpest result came from a question I nearly recorded as empty. Ask ChatGPT who to buy off-plan from in Dubai and it names developers, not brokerages, in 5 of 5 runs. Emaar appears in all 5 and is the first name listed in all 5. Not one brokerage is named, so a brokerage optimising for that question is competing for a slot the engine doesn't appear to offer. That section is further down, with the answers quoted in full.
Collected 16 September 2026. Requests issued as one concurrent batch at 02:43:13Z; responses returned between 02:43:27Z and 02:46:51Z, a span of 3 minutes 24 seconds
Engine: ChatGPT search mode, read through the DataForSEO LLM Scraper (live/advanced)
10 questions, 5 runs each, 50 runs total
Location parameter: United Arab Emirates (2784). The endpoint has no sub-national entry, so Dubai is named in the question text
Language: English. force_web_search requested (the vendor doesn't guarantee search data appears)
Model: not reported by the endpoint. search_results and fan_out_queries came back null in all 50 runs
No byte-identical duplicate responses: 50 distinct response hashes across 50 runs. This rules out simple byte-for-byte caching, not caching that recomposes or partly reuses an earlier answer
Register: DLD broker export, downloaded 16 September 2026, 43,377 broker card rows covering current and expired cards, yielding 13,601 office registration numbers that resolve to 13,485 distinct entities after branch merging. Names are matched against a derived firm-level file, firms.csv, SHA-256 1b4546fbbd28340af3208fd2e7b19b38fda1b1ec9cf26716dab4b38d2d4882ae. The raw export's own SHA-256 and the command that rebuilds firms.csv from it are in the bundle manifest. Dubai Pulse updates the dataset, so a later download will match neither hash
Extraction: recall 90.9% on registered brokerages, 59.4% on all named organizations, precision 85.9%, against a 12-run ground truth read blind to the extractor's output. Scored by a rule and a script both committed before the scorer ran
Map-block parse fidelity: 561 of 561 entries stored, zero drift, all 50 runs
Pack: dubai-brokerage-v1, prompt strings in the appendix. Pack hash fda99b65dcec970bf572e1c3c54630e44a1b838fe34af4cd2f9447ecdcea7359, recorded in the ledger against all 50 runs. The pack file was committed as 330df4a2e02ae77315faf40b3ff658d367e25465 on 16 September 2026 at 10:43:08 (UTC+8), five seconds before the first request went out
All 50 raw responses stored
The pack hash is stored on every one of the 50 run records, and it reproduces from the pack file in the repository, so the prompt strings can be shown not to have changed during or after collection. That much is checkable from the bundle without taking my word for anything.
What the commit does not do is pre-register the questions. It lands five seconds before the first request, so committing the pack and running it were one scripted action. It shows the pack was not edited afterwards to match whatever came back. It cannot show the questions were fixed before I had any idea what the answers would look like, and nothing else here shows that either. Edition 2 commits and pushes the pack on a different day from collection, which is the only version of this claim worth making.
The register snapshot was committed after collection, which is the right order: the questions went out first, and the names that came back were matched against the register afterwards. The analysis outputs behind every figure on this page were generated at 18:34 on 17 September, after the last change to that snapshot, so the file whose hash appears above is the one the numbers were produced against.
What this study does not claim
The register snapshot contains 13,601 office registration numbers, drawn from 43,377 broker card rows covering both current and expired cards. Office registration numbers are the entity-level reference set this study matches against. This edition named 142 distinct organizations once not-a-firm labels are excluded. 89 of those resolved to a registration number, and 42 of the 89 appeared in 5 or more runs. Put 42 over 13,601 and you get "99% of Dubai brokerages are invisible to AI."
That number is worthless and I'm not publishing it. ChatGPT returns a handful of names per answer no matter how many firms exist, so the invisible share is fixed by the format of the answer, not by anything the firms did. Any market, any register, any denominator gives you 99-point-something. It's the same shape as observing that 99% of restaurants aren't on a top-10 list.
What the slot count does tell you is worth having: across 50 runs of 10 different questions, the engine named 89 distinct organizations, 42 of them more than four times. That's the size of the pool actually in play. Everything below is about who is in it and how steady their place in it is.
Merging spelling variants moves individual pairs, not the middle
This section comes first because it's the part I'd argue with hardest if someone else had written it.
ChatGPT calls the same firm several things. Betterhomes appears as "Betterhomes", "betterhomes - Dubai (Head Office)" and "betterhomes - Marina Office". fäm Properties holds 14 office registration numbers in the DLD register. Count those as different firms and the answers look less consistent than they are.
So I resolved every name against the register, merged branches and spelling variants onto one licence identity, and recomputed.
| Raw strings | Resolved | |
|---|---|---|
| Median, all pairs | 0.304 | 0.318 |
| Mean, all pairs | 0.312 | 0.335 |
| Pairs that changed | 45 of 99 (39 up, 6 down) | |
Nearly half the pairs moved and the middle barely did. That's worth separating, because the two facts point different ways. Resolution matters a great deal for any individual firm's number and very little for the headline.
The strict counts move more. On "most reputable", the number of firms named in all 5 runs goes from 1 to 4 once variants are merged. The 4 are Allsopp & Allsopp, Betterhomes, haus & haus and Metropolitan Premium Properties. Betterhomes was in every run the whole time, under three different strings. Anyone measuring this without a register would have reported 1.
I had drafted a version of this piece claiming that a large share of the volatility other people report is their own name-matching. The table above says that's wrong at the aggregate level, so the claim is gone. It survives only for per-firm figures, where it's severe.
How much two answers to the same question overlap
I took the set of firms named in each run and compared every pair inside a question, using Jaccard: firms named in both runs divided by all firms named across the two. A score of 1.0 means the two runs named exactly the same firms. A score of 0 means nothing in common.
10 questions at 5 runs each gives 100 pairs. One pair is undefined, because both of its runs were scored as naming nobody and Jaccard has no value when both sets are empty. That leaves 99 scorable pairs, and both this section and the resolution table above use the same 99.
Jaccard is stricter than it sounds. For two lists of similar length, 0.32 means about 48% of each list also appears in the other, and 0.66 means about 80%.
| Question | Median overlap | Lowest pair | Highest pair | Named in all 5 | Distinct firms |
|---|---|---|---|---|---|
| Buy in Dubai Marina | 0.66 | 0.56 | 0.82 | 10 | 21 |
| Rent an apartment | 0.50 | 0.32 | 0.61 | 6 | 27 |
| List an apartment to sell | 0.39 | 0.26 | 0.71 | 6 | 30 |
| Luxury property | 0.37 | 0.19 | 0.48 | 4 | 38 |
| Best for investors | 0.29 | 0.08 | 0.73 | 2 | 34 |
| Best agencies | 0.28 | 0.12 | 0.79 | 2 | 39 |
| Most reputable | 0.28 | 0.19 | 0.59 | 4 | 31 |
| Buy an apartment | 0.24 | 0.10 | 0.56 | 1 | 42 |
| Relocating | 0.13 | 0.00 | 0.33 | 0 | 22 |
| Off-plan | 0.00 | 0.00 | 0.06 | 0 | 33 |
Across all 99 scorable pairs the median is 0.32 and the mean is 0.34. The median of the 10 question medians is 0.29. I'm quoting the pair-level figure throughout; both are in the data.
The min and max columns matter more than I expected. "Best agencies" runs from 0.12 to 0.79. The same question, minutes apart, produced one pair of answers that barely agreed and another that nearly matched. A single median hides that.
Two limits on how hard to read these medians. Each rests on 10 pairs drawn from only 5 runs, so the pairs share runs and aren't independent of each other. And the off-plan row is my extractor's number rather than the engine's behaviour, for reasons in the next section but one.
Dubai Marina
Naming a district instead of the city produced the steadiest answers in the edition: median 0.66, and 10 firms named in every one of the 5 runs out of 21 total.
Against the matched question ("I want to buy an apartment in Dubai", 0.24) that's 2.75x. Against the edition median, 2.1x. Pick whichever you prefer, and note this rests on one question in one edition. n = 1. I'm not going to defend it further until Singapore.
My first guess was that Marina has fewer agencies to choose from. My second was that a district name triggers the map block, which is fed by a places provider and ought to be steadier than prose. The data doesn't support the second one cleanly. Marina is the most stable question on map-block names alone (0.68) and also on prose names alone (0.54), and it's the highest of all 10 questions on both. The Marina result isn't explained by the map-block names alone.
The off-plan question returned developers, not brokerages
This is where my first pass was wrong, and the correction is the more interesting result.
I had recorded 2 of the 5 off-plan runs as returning no firms. They weren't empty. My extractor reads the map block, bulleted lists and tables, and those 2 runs answered in a single bold comma-separated sentence, which no channel picks up. Here's one of them in full:
For a Dubai off-plan purchase, I'd shortlist Emaar, Meraas, Nakheel, Dubai Properties, Ellington, and Majid Al Futtaim and then choose based on the specific project, not simply the developer's name. These are among the developers participating in Dubai's official first-time-buyer programme.
Before paying a booking deposit, verify with DLD/RERA that the project is registered, has its own escrow account, has the required permits, and that the developer is registered.
Read by hand across all 5 runs, the off-plan answers are the most consistent in the edition. Emaar appears in 5 of 5 and is the first name listed in all 5. Nakheel 5 of 5, Meraas 4 of 5, Sobha 4 of 5, Ellington 3 of 5.
Not one of them is a brokerage. Ask who to buy off-plan from and ChatGPT answers with developers, reliably, and then tells you to verify escrow and permits with DLD before paying a deposit. A brokerage optimising for that question is competing for a slot the engine doesn't appear to offer.
The 0.00 in the overlap table is my extractor's number, not the engine's behaviour. I've left it in the table and flagged it here rather than quietly fixing it, because the gap is real and it affects any other run where a firm is named mid-sentence. The same correction changes the position-1 counts further down, and I've marked where.
The relocating question asked me where I was
One run of the relocating question was 142 characters and named nobody:
I can help compare Dubai estate agents, but I need your location confirmation first because the available location signal is only an estimate.
That's the engine declining to answer until it knows where I am. It's the only genuinely empty run in the edition, and it lines up with a limit of the method: the scraper's location list has no entry below United Arab Emirates, so the geographic signal reaching ChatGPT is weaker than a real Dubai buyer's would be.
Every Dubai brokerage named in at least 5 of the 50 runs
42 registered firms cleared that bar. "Registered" here means the name resolved to an office registration number in the 16 September snapshot. It does not mean the firm is trading: a licence can sit on the register with no active brokers attached, and at least one entry below is in that position. Each one is matched to the DLD broker register; the register's own spelling is in the second column, and branch listings and spelling variants are merged onto one licence identity.
Read it as reach, not as a ranking. Rows are ordered by how many runs named the firm, because that is the readable order, and the ordering is not a ranking of brokerages. These figures pool 10 different questions, and renting, luxury and listing-to-sell should produce different firms even from a perfectly steady engine.
Read the interval, not the percentage. Every rate here comes from 50 runs, which is a small sample. The interval column gives descriptive 95% binomial intervals, computed as though the 50 runs were independent trials. They are not: 5 runs belong to each question, the questions are deliberately different, and a firm's chance of being named is correlated within a question. So read these as a rough width on a small sample rather than as confidence intervals for how ChatGPT will behave in general. Almost every pair of firms in this table has overlapping intervals, and treating the order of these rows as a ranking is the specific mistake this column exists to stop.
| Firm | Registered as (DLD) | Runs | % of 50 | 95% binomial | Questions | First |
|---|---|---|---|---|---|---|
| Exclusive Links | EXCLUSIVE LINKS REAL ESTATE BROKERS | 35 | 70% | 56–81% | 9 | 1 |
| fäm Properties | F A M REAL ESTATE BROKER L.L.C | 22 | 44% | 31–58% | 6 | 10 |
| Betterhomes | BETTER HOMES (L.L.C) | 20 | 40% | 28–54% | 7 | 8 |
| Harbor Real Estate | HARBOR REAL ESTATE BROKER L.L.C | 20 | 40% | 28–54% | 6 | |
| White & Co | WHITE AND WHITE REAL ESTATE L.L.C | 20 | 40% | 28–54% | 5 | 4 |
| Allsopp & Allsopp | ALLSOPP & ALLSOPP REAL ESTATE BROKER L.L.C | 19 | 38% | 26–52% | 7 | |
| Driven Properties | DRIVEN PROPERTIES L.L.C | 19 | 38% | 26–52% | 7 | 2 |
| District Real Estate | DISTRICT REAL ESTATE L.L.C | 17 | 34% | 22–48% | 7 | |
| 99 Real Estate Brokers | 99 REAL ESTATE BROKER (L.L.C) | 15 | 30% | 19–44% | 5 | |
| Metropolitan Premium Properties | METROPOLITAN PREMIUM PROPERTIES L.L.C | 15 | 30% | 19–44% | 5 | 2 |
| Christie's International Real Estate Dubai | CHRISTIES INTERNATIONAL REAL ESTATE L.L.C | 14 | 28% | 17–42% | 5 | 4 |
| D&B Properties | D A N D B PROPERTIES | 13 | 26% | 16–40% | 5 | 5 |
| haus & haus | HAUS & HAUS REAL ESTATE BROKER L.L.C | 13 | 26% | 16–40% | 4 | 3 |
| Real Choice Real Estate | REAL CHOICE REAL ESTATE BROKERS (L.L.C) | 13 | 26% | 16–40% | 5 | |
| Springfield Properties | SPRING FIELD PROPERTIES L.L.C S.O.C | 13 | 26% | 16–40% | 6 | |
| Next Level Real Estate | NEXT LEVEL REAL ESTATE | 12 | 24% | 14–37% | 6 | |
| Dacha Real Estate | DACHA REAL ESTATE | 11 | 22% | 13–35% | 5 | 3 |
| Lovisa Homes | LOVISA HOMES REAL ESTATE BROKER L.L.C | 11 | 22% | 13–35% | 4 | |
| Knight Frank | KNIGHT FRANK REAL ESTATE BROKERAGE L.L.C | 10 | 20% | 11–33% | 3 | |
| Provident Real Estate | PROVIDENT REAL ESTATE BROKER L.L.C | 10 | 20% | 11–33% | 3 | |
| Engel & Völkers Dubai | EV REAL ESTATE BROKERAGE L.L.C | 9 | 18% | 10–31% | 3 | |
| Xsite Real Estate | XSITE REAL ESTATE BROKERS L.L.C | 9 | 18% | 10–31% | 4 | |
| BSO Real Estate Management | B S O REAL ESTATE MANAGEMENT L.L.C S.O.C | 8 | 16% | 8–29% | 4 | |
| McCone Properties | MCCONE PROPERTIES | 8 | 16% | 8–29% | 5 | 1 |
| Al Ayan Real Estate | ALAYAN REALESTATE BROKER | 7 | 14% | 7–26% | 3 | |
| Elite Property Brokerage | E L I T E PROPERTY BROKERAGE L.L.C | 6 | 12% | 6–24% | 4 | |
| House Hunters | HOUSE HUNTERS REAL ESTATE BROKERS L.L.C | 6 | 12% | 6–24% | 3 | |
| Impressive Real Estate | Impressive Real Estate | 6 | 12% | 6–24% | 4 | |
| Kharz Real Estate | KHARZ REAL ESTATE L.L.C | 6 | 12% | 6–24% | 4 | |
| Luxury Homes Real Estate | LUXURY HOMES REAL ESTATE L.L.C | 6 | 12% | 6–24% | 2 | |
| Luxury One Real Estate | LUXURY ONE REAL ESTATE L.L.C | 6 | 12% | 6–24% | 2 | |
| The Luxury Real Estate | THE LUXURY REAL ESTATE BROKER L.L.C | 6 | 12% | 6–24% | 2 | |
| Unique Properties | UNIQUE PROPERTIES L.L.C | 6 | 12% | 6–24% | 4 | |
| Aaronz & Co | AARONZ & CO REAL ESTATE L.L.C | 5 | 10% | 4–21% | 1 | |
| Agency Eight | AGENCY EIGHT REAL ESTATE L.L.C | 5 | 10% | 4–21% | 1 | |
| Binayah Properties | BINAYAH PROPERTIES L.L.C | 5 | 10% | 4–21% | 2 | |
| Bluechip Real Estate | BLUECHIP REAL ESTATE BROKER l.l.c | 5 | 10% | 4–21% | 3 | |
| Espace Real Estate | ESPACE REAL ESTATE BROKER L.L.C | 5 | 10% | 4–21% | 1 | |
| Range International Property Investment | RANGE INTERNATIONAL PROPERTY INVESTMENT ONE PERSON COMPANY L.L.C | 5 | 10% | 4–21% | 1 | |
| Seven Luxury Real Estate | SEVEN LUXURY REAL ESTATE L.L.C | 5 | 10% | 4–21% | 1 | |
| TheRealtorDubai | THEREALTORDUBAI REAL ESTATE BROKERAGE L.L.C | 5 | 10% | 4–21% | 2 | |
| Union Square House | UNION SQUARE HOUSE REAL ESTATE BROKER L.L.C | 5 | 10% | 4–21% | 4 |
The register writes almost every firm in capitals. "Impressive Real Estate" is the exception in this table and is stored in the DLD export exactly as shown, mixed case. I've reproduced each spelling as the export has it rather than tidying it.
These counts are floors. The audit found that firms named in a sentence rather than a list are missed, so any firm recommended in prose is undercounted here by an unknown amount. Chestertons at 3 runs, Provident at 10 and Dacha at 11 each have at least one prose mention the extractor did not catch. The audit covered 12 of the 50 runs, so I can't say how far this moves the rest of the table.
37 of the 42 appeared in fewer than 4 answers in 10. Only 5 reached 20 runs or more, and only Exclusive Links passed 30. Reach is out of all 50 runs; on the 49 that named anyone, Exclusive Links is 71% rather than 70%.
The "Questions" column is the one I'd watch if I ran a brokerage, because it captures something raw run count doesn't: whether a firm shows up across different kinds of buyer, or repeatedly under one. 10 questions is itself a small sample of the questions people ask. Exclusive Links turned up under 9 of the 10 questions. Espace, Agency Eight and Aaronz & Co each appeared under 1, which means their presence rests entirely on one kind of buyer. Two firms can share a reach percentage and a confidence interval and still be in completely different positions on this column.
23 more licensed firms, named in 2 to 4 runs
No rate published for these. At 2 to 4 appearances out of 50 the interval runs from about 1% to 19% and the number says more about the sample than the firm. My threshold for a published rate is 5. They're listed because they were named and they're licensed, which is checkable either way.
| Firm | Registered as (DLD) | Runs |
|---|---|---|
| DSQ Real Estate | D S Q REAL ESTATE | 4 |
| Lone Wolf Real Estate | LONE WOLF REAL ESTATE | 4 |
| Luxury Concierge Real Estate | LUXURY CONCIERGE REAL ESTATE BROKER L.L.C | 4 |
| Smart Folks Real Estate | SMART FOLKS REAL ESTATE BROKERS L.L.C | 4 |
| Strada Real Estate | STRADA REAL ESTATE BROKERAGE L.L.C | 4 |
| Abode Property | ABODE PROPERTY L.L.C | 3 |
| Algebra Real Estate | ALGEBRA REAL ESTATE BROKER LLC | 3 |
| Buy Sell Let Real Estate | BUY SELL LET REAL ESTATE | 3 |
| Chestertons | CHESTERTON INTERNATIONAL REAL ESTATE BROKERAGE L.L.C | 3 |
| Dream Home Real Estate | DREAM HOME REAL ESTATE BROKER L.L.C | 3 |
| Homes 4 Life | HOMES 4 LIFE REAL ESTATE BROKER (L.L.C) | 3 |
| LuxuryProperty.com | LUXURY PROPERTY L.L.C | 3 |
| Prima Luxury Real Estate | PRIMA LUXURY REAL ESTATE BROKER L.L.C | 3 |
| Wikihomes | WIKI HOMES REAL ESTATE BROKER L.L.C | 3 |
| ALH Properties | A L H PROPERTIES L.L.C | 2 |
| Aeon & Trisl | AEON & TRISL REAL ESTATE BROKER L.L.C | 2 |
| EQT Real Estate | EQT REAL ESTATE L.L.C | 2 |
| Huaxia Real Estate | HUAXIA REAL ESTATE BROKER L.L.C | 2 |
| Powerhouse | POWER HOUSE PROPERTIES BROKERS | 2 |
| Prestige Luxury Real Estate | PRESTIGE LUXURY REAL ESTATE | 2 |
| Dubai Sotheby's International Realty | SIR REAL ESTATE L.L.C | 2 |
| Takween Aldar | TAKWEEN ALDAR REAL ESTATE L.L.C | 2 |
| The Property | THE PROPERTY REAL ESTATE | 2 |
A further 24 licensed firms were named exactly once. I've left those out.
21 names I couldn't match to the register
Four are my own matching falling short, not missing firms.
| As named | What it is | Runs |
|---|---|---|
| fäm Living | a fäm Properties sub-brand, absent from the register | 5 |
| Provident Estate | Provident Real Estate Broker (already above at 10 runs) | 3 |
| Luxhabitat Sotheby's International Realty | Luxhabitat is a marketplace; the brokerage SIR Real Estate appears above | 2 |
| White & Co | a second string form of White and White Real Estate; not counted in that firm's 20 runs | 1 |
The other 17 returned nothing against the DLD broker export under any spelling I tried. Trade names diverge from brands constantly: White & Co is registered as White and White Real Estate, D&B Properties as D A N D B Properties. Others may be licensed in another emirate, since my location parameter covers the whole UAE. Some may be international firms with no Dubai office. I haven't checked the DLD developer register or the free-zone register against these, so I'm not calling any of them unlicensed or non-existent.
| As named by ChatGPT | Runs |
|---|---|
| Alexander Kaminsky Real Estate | 5 |
| First Class Property Management | 4 |
| Sacred Lands Group | 4 |
| Dubai Investissement | 3 |
| Excel Properties | 3 |
| Marina Immo Dubai | 2 |
| One Investments | 2 |
| PH Luxury: Real Estate Agents & Brokerage Firm in Dubai | 2 |
| Pacific Real Estate Investment LLC | 2 |
| Abbasi One Real Estate Dubai | 1 |
| Arabian Estates LLC | 1 |
| Dubai Luxury Homes | 1 |
| DubaiBulls | 1 |
| Mayfair Luxury Real Estate | 1 |
| Real estate agency UAE Assets | 1 |
| Rosenheim Luxury Properties | 1 |
| Salim Real Estate-Buy,Sell,Rent a Property Monthly in Dubai | 1 |
Which firm opens the answer, question by question
Under the extractor's reading, 47 of the 50 runs opened with a named firm. The off-plan correction above moves that to at least 49: both prose runs open with Emaar, which the extractor didn't see. The table below is the extractor's version, with the affected questions marked, because recomputing it by hand for two runs and not the other 48 would mix two methods in one column.
| Question | Distinct firms at position 1, across 5 runs |
|---|---|
| Best agencies | 2 |
| Luxury property | 2 |
| Rent an apartment | 2 |
| Off-plan (3 of 5 runs scored) | 2 |
| Buy in Dubai Marina | 3 |
| List an apartment to sell | 3 |
| Relocating (4 of 5 runs scored) | 3 |
| Buy an apartment | 4 |
| Best for investors | 4 |
| Most reputable | 4 |
Pooled across the edition, 14 different firms held the opening slot, but pooling 10 questions overstates the churn, so the per-question cut above is the honest one. First place varies across runs, within bounds: two questions held the top slot with 2 firms across 5 runs, and no question used more than 4.
The most repeated single holder is fäm Properties, first in 4 of 5 runs of "best agencies" and 10 times across the edition. Of the 47 scored openings, 44 are registered brokerages and 3 are developers, all 3 on off-plan. Counting the 2 prose runs the extractor missed, developers hold the opening slot in all 5 off-plan runs.
The First column in the reach table sums to 43, not 44. The missing opening belongs to LuxuryProperty.com (DLD: LUXURY PROPERTY L.L.C), which opened one answer but was named in only 3 runs, so it sits in the 2-to-4-run table rather than the one above. 43 plus that one, plus the 3 developer openings, accounts for all 47 scored runs.
What this method can't see
Firms named in a sentence are missed. My extractor reads the map block, bulleted lists and tables, and nothing else. A recommendation written as prose is invisible to it.
The audit put a number on that. 50 of the 54 misses sit in inline prose, across four different questions. It's one failure mode rather than a scatter of unrelated gaps, and it's the same one that had me record the off-plan answers as empty.
What the extraction audit found
An earlier version of this audit reported 93.8% recall. That figure was wrong in a way worth describing, because the mistake is easy to make. 98 of its 227 names had been pulled by regex out of the map block rather than read by anyone, so 43% of the audit was comparing the extractor against itself. It also happened to sample none of the runs where the extractor had failed.
Redone on 12 runs, with ground truth written by a separate model instance that never saw the extractor's output. Blind to the extractor, not independent of me: I wrote both the extractor and the ruleset the reader worked from, and the detail is below the table.
| Measure | Count | Value | 95% binomial |
|---|---|---|---|
| Recall, registered brokerages only | 70 of 77 | 90.9% | 82.4 to 95.5 |
| Recall, all named organizations | 79 of 133 | 59.4% | 50.9 to 67.4 |
| Precision, checklist labels included | 79 of 92 | 85.9% | 77.3 to 91.6 |
| Precision, checklist labels excluded | 79 of 81 | 97.5% | 91.4 to 99.3 |
| Map-block parse fidelity, all 50 runs | 561 of 561 | 100% |
The two recall figures answer different questions, and the gap between them is almost entirely regulators, portals and named systems: DLD, RERA, Trakheesi, Oqood, Dubai REST, Google, Bayut, Property Finder, TruBroker. Those get named constantly and none of them belongs in a brokerage scoreboard. 90.9% is the figure that bears on the tables above. 59.4% is what the extractor does with everything else.
Both interval columns on this page give descriptive 95% binomial intervals, calculated as though each run or each name were an independent trial. Runs are clustered by question and names are clustered within runs, so these are a rough width on a small sample, not population-level confidence intervals for how ChatGPT behaves in general.
The two precision figures are worth separating. Excluding the checklist labels the extractor is right 79 times out of 81, and the 13 errors it does make are almost all one known class: words like "Budget" and "Payment plan" that open a numbered list and get read as names. That class is visible, countable and in the published data.
All of these are scored against the answer text, which is what the ground-truth reader saw. Scoring the map block as well changes recall by nothing at all, because every map entry the reader saw was repeated in prose, and it drops apparent precision to 34.2% purely by adding names the audit never covered. That lower number measures the audit's scope rather than the extractor, so it is not quoted as precision here. Both variants are in the published scoring output.
Two scorers were written for this, independently, and they were compared after both had run. They agree on every one of the 133 hand-read names and all 92 extracted names, and disagree on 3 rows, all about which names count as brokerages rather than about whether the extractor found them. One scorer counted Sobha Realty, which holds a licence but is a developer everywhere else in this study. The other counted two copies of a map-listing title as a firm. That disagreement is the whole gap between 90.9% and 91.0%. The figure quoted here is the one the committed rule produced, including where the rule is arguably worse than the alternative, and edition 2 fixes both defects in the rule rather than in the number.
Parse fidelity is a code test rather than a hand read, and it's reported separately for that reason. Every map-block entry in every run became a stored mention, with no drift in position, title or URL.
All 7 missed brokerages were named in prose. Chestertons, Dacha and Provident in one sentence of the Dubai Marina question ("A Dubai Marina directory also currently identifies Chestertons, Dacha, Espace and Provident"). Allsopp & Allsopp and Betterhomes in one sentence of the investor question. Sotheby's and fäm the same way.
Three runs were audited separately because the extractor had scored them as empty: both off-plan runs and the clarifying answer. Recall there was 0 of 14. Those numbers are reported on their own and never folded into the 12-run figures, since deliberately sampling known failures would bias the result.
Ground truth came from a separate model instance with no access to the extractor's output, one session per run, working from a fixed ruleset. The ruleset changed 6 times over the first 6 runs and held for the last 9; 3 runs were re-read under the final version, 1 run was replaced after I supplied a name list before it ran, and 4 results were amended by me afterwards, each time to add a name the ruleset covered and the session had left out. I wrote the extractor and the ruleset, so this is a blind read rather than an independent one.
Extractor false positives are still in the counts. They're checklist labels like "Budget" and "Payment plan" that open numbered lists, and the resolver classes them as not-a-firm rather than dropping them. This edition has 170 distinct names, 83 of which appear in exactly one run. 28 of the 170 fall in the not-a-firm class, and 27 of those 28 are singletons. Excluding them gives 142 distinct names and 56 singletons. None reached position 1 in any run.
No control question. All 10 prompts ask for an agency, so this edition has no null condition, and precision (85.9% with checklist labels counted) is the only bound on the extractor over-firing. Edition 2 adds a question that asks about the buying process and names no firm; unsolicited naming there will be worth more than anything in the tables above.
The model is not identified. The endpoint did not report a model identifier on any of the 50 runs. This measures what came back through one vendor's ChatGPT search-mode interface on 16 September 2026. It cannot be attributed to a named model, and if the model behind that interface changed the day after, nothing here would show it.
One moment, not many. All 50 runs returned inside 3 minutes 24 seconds of each other. That limits how much ordinary market change could plausibly explain the differences between the first answer and the last. It does not identify what produced them: the endpoint returned no search_results and no fan_out_queries, so the retrieval behind each answer is not visible here and the variation cannot be attributed to any particular mechanism. The flip side is that this measures instability within one moment and says nothing about whether the same question answered next week returns the same firms. Those are different quantities and this edition only has the first. Edition 2 repeats the pack on separate days to get the second.
Country-level location. Dubai reaches the engine through the question text, not the geography setting, and one run stopped to ask me where I was. This isn't the same as asking from a Dubai IP, and I haven't tested whether the two agree. That calibration run is the first thing scheduled for edition 2.
English only. Arabic and Russian are real buyer languages in this market and transliterated firm names are a matching problem this method doesn't handle.
One engine, one day. ChatGPT search mode on 16 September 2026, 5 runs per question. Gemini, Perplexity and Google's AI Overviews may behave completely differently.
Corrections
This log is permanent and carries forward into every edition. Each entry is something published or drafted here that turned out to be wrong.
- Recall of 93.8% was not a real measurement. 43% of that audit's ground truth had been generated by the same regex the extractor uses, so it was scoring itself. Replaced by a blind 12-run read, rescored under a written rule (entry 6).
- Two off-plan runs recorded as empty were not empty. They named six developers each in a prose sentence the extractor doesn't read. The off-plan overlap figure of 0.00 is an artefact of my code, and it's still in the table above, labelled.
- "Zero names common to all five runs of the most-reputable question" was wrong. It was 1 on raw strings and 4 once names were resolved to licences. The claim was measuring my matching, not the engine, and it was cut before publication.
- The framing that AI volatility is mostly other people's entity-resolution failure did not survive its own test. Resolution moved 45 of 99 pairs and the median by 0.014. The claim now holds only for individual firms' figures.
- Recall of 91.0% went into a near-final draft with no script behind it. It could not be reproduced from the repository, because the scoring had been done in conversation and never written down. The rule was then written, committed, and run once, and the figure came out at 90.9% rather than 91.0%. The two differ over 3 classification rows, not over anything the extractor did. Precision of 85.9% had the same problem and reproduced exactly. The numbers were close to right; the process that produced them was not, and this is the second recall figure on this page to have had that problem. The first is entry 1.
- A draft of this page stated 15 extractor false positives and then subtracted 28. The correct count is 28 not-a-firm names, 27 of them singletons. The 142 and 56 figures were right; the label on them was not. Caught in review, before publication.
What was already known
SparkToro and Gumshoe ran 12 recommendation prompts 60 to 100 times each across ChatGPT, Claude and Google's AI during November and December 2025, 2,961 runs in total, published January 2026. For ChatGPT and Google's AI in Search specifically, they put the odds of two runs of one prompt returning the same list of brands at under 1 in 100, and the same list in the same order at under 1 in 1,000.
Their categories were consumer ones: chef's knives, headphones, cancer care hospitals, marketing consultants, science fiction novels. In their headphone set, four brands appeared in 55% to 77% of 994 responses to prompts written by different volunteers.
My edition replicates the core instability result on a market they didn't cover, with a different design: one frozen string per question rather than volunteer-written variants, and a licensed-firm register underneath. I'm not going to line Exclusive Links' 70% up against their 55–77% band, because the two numbers are built differently. Theirs pools many phrasings of one consumer category; mine pools 10 fixed questions in one city. They aren't the same measurement and treating them as one would be the kind of borrowed-number comparison I'd flag in someone else's study.
The part of this I think is most useful to build on is resolving every named firm against a government licence register. That's what makes the Betterhomes case visible. I haven't run a literature search, so treat that as a view about what's worth doing next rather than a claim about who has done it before.
What I'd take from this if I ran a brokerage
Being named once tells you close to nothing. A firm can appear in one answer and be gone from the next, and both are normal.
The number worth tracking is how often you appear across many runs of many questions, and how many different questions reach you. A firm at 70% across 9 questions and a firm at 10% under 1 question are in different positions even if both can produce a screenshot.
Which is also the problem with a screenshot. "Best agencies" produced one pair of runs overlapping at 0.79 and another at 0.12, minutes apart, on the same question. Either one makes a convincing image.
And check what question you're competing for before optimising for it. Off-plan returns developers in 5 of 5 runs.
Edition 2 runs 18 questions on a frozen prompt set at deeper repetition, with a control question added, recommendation strength scored rather than presence counted, and a same-day calibration run from a Dubai IP. Singapore follows on the same pack.
Disclosure
We sell AI visibility services to real estate businesses, so we have a commercial interest in this subject being taken seriously. Brokerages named here are potential customers of ours. No brokerage named in this study is a client, and if one becomes a client we'll mark the row in the edition where that happens.
We're not a Dubai brokerage and don't compete with any firm in these tables. We hold no commercial relationship with the Dubai Land Department, Bayut or Property Finder.
The prompt set and the register snapshot were fixed and recorded before collection. The brand-to-licence bridge could only be built after reading the names, so those decisions came after the results and are recorded with their basis. Method changes are versioned and disclosed per edition.
Check this yourself
Everything needed to recompute every number above is published, not available on request:
- All 50 raw API responses, unmodified, as stored at collection
- The frozen prompt pack, with the commit hash and date it was frozen
- The SHA-256 of the DLD export and of the derived firm-level file, with the command that rebuilds one from the other. The export itself is linked rather than republished
- The run-level extraction CSV: run, question, run index, firm, matched form, position
- The brand-to-licence bridge, with a reason recorded per entry
- The matcher source, and the 12-run ground truth the recall and precision figures were scored against
The bundle is at realestateseo.ph/research/dubai-01/data/. The zip is 332,515 bytes, SHA-256 e9754176990a38372f44a088dacb6a377dbc49d0c3db85aab5bbd339b72db54e. The manifest sits outside the zip as plain text, listing a SHA-256 for every file, so the three hashes printed in the stamp above can be checked without downloading anything.
The manifest also lists what is deliberately left out and why: the DLD export itself, which is linked and hashed rather than republished, and the working notes from the run, which contain claims this page later withdrew.
Appendix: the 10 questions, verbatim
Every run sent one of these strings exactly as written, with no system prompt and no conversation history.
2. I want to buy an apartment in Dubai. Which real estate agency should I use?
3. Which real estate agency should I use to buy a property in Dubai Marina?
4. Which real estate agency in Dubai is best for property investors?
5. Which agency handles luxury property in Dubai?
6. Which real estate brokerages in Dubai are the most reputable?
7. Who should I buy an off-plan property from in Dubai?
8. I am relocating to Dubai. Which estate agent should I contact?
9. Which agency should I use to rent an apartment in Dubai?
10. Which agency should I list my Dubai apartment with to sell it?