Research · AI takeoff accuracy · Sources checked 2026-08-19

AI takeoff accuracy: what vendors advertise, and what has actually been measured

Three AI takeoff vendors put a number in the high nineties on their sites. We went looking for where those numbers come from, and for anyone who has measured one independently. This is everything we found, with the sources, and why Bildrix publishes no percentage at all.

SHORT ANSWER

How accurate is AI takeoff software? Nobody outside the vendors has measured it in a way you could repeat. The advertised figures (98%, 94.5 to 98.5%, 99%+) come with no named plan set, no definition of a correct detection, and no error bars. The one peer-reviewed study found a single first-time user’s corrected Togal takeoff within 5% of a manual On-Screen Takeoff count on a few small floor plans, and concluded that relying on the AI output alone is not advisable. Field reports put the raw first pass anywhere from about 60% to the mid-nineties, depending on the sheet.

The honest answer is that accuracy belongs to the tool and the drawing together, and the useful question is not what a tool averages, but whether it tells you where it is unsure.

8 vendors checked on their own sites3 publish a headline percentage1 peer-reviewed study · 1 trade-press testPublic benchmark datasets found: 0
01 · Advertised

What the vendors say on their own pages.

Read on each vendor’s own site on 2026-08-19. Quotation marks mean verbatim; nothing here is relayed from a third party. Three of the eight publish a headline figure; one discloses a distribution in a recorded webinar; four publish no number.

Table 1 · Advertised accuracy claims, AI construction takeoff tools · checked 2026-08-19
VendorAdvertisedWhere it appearsMethod disclosedSrc
Togal.AI98%“up to” · “on floor plans”Homepage hero puts it at “up to 98%”; the benefit block headed “Increase Accuracy” drops the “up to”, alongside “5x faster” and “Takeoff in Minutes. Not Days.” The pricing FAQ calls it “the fastest, most accurate estimating solution on the market.”None. No footnote, no dataset, no definition of what counts as a correct detection.[1][2]
Kreo94.5 to 98.5%“Expert quality”On the agentic-workflow page: “Expert quality of 94.5-98.5%”, next to “up to 10,000 drawings per minute”. The AI takeoff product page itself carries no figure.“Trained on thousands of projects.” No dataset, no definition, no error bars.[3][4]
Beam AI99%+QA-checked, 24 to 72 hoursHomepage: “99%+ accurate, QA-checked takeoffs and estimates in 24-72 hours”. Its Kreo comparison page sells “100% automated takeoffs” and, in the same table, states “Every takeoff is reviewed by a human QA team.”None published. The figure describes a human-reviewed deliverable, not an unreviewed AI pass: a legitimate thing to sell, and not comparable to a real-time detector's claim.[5][6]
Drawer AINo figure on sitedistribution in a recorded webinarIn the posted demo-webinar recording, the presenter gives results over a 250-project databank: about 15% of projects at 100%, 60% at 95% or better, 90% at 85% or better, and puts the point where ROI “kicks in” north of 80. Vector PDFs only; scanned sets do not run.Self-measured on a databank that is not published, but it is a distribution with the failures in it, which no headline number is.[7]
BobyardNo figuretime claim onlyFAQ: takeoff time cut “50-90%, depending on drawing quality, trade, and the estimator review process.”Not applicable[8]
STACK (Floor Plan AI)No figureThe product page promises “speed and accuracy”; there is no number anywhere on it.Not applicable[9]
Takeoff Boost (On Center / PlanSwift)No figure“Highly Accurate AI… trained on millions of construction projects.” PlanSwift's own FAQ is plainer: Takeoff Boost is “a head start, not a finished bid.”Not applicable[10][11]
Houzz Pro AutoMateNo figurescope limit disclosedHelp center: AutoMate runs on “simplified floor plans and elevations only at this time.”Not applicable[12]
02 · Measured

Everything that has actually been measured.

Three kinds of evidence exist: one peer-reviewed study, one trade-press test, and field reports. Each is smaller than it looks from its headline.

PEER-REVIEWED · 1 STUDY

One peer-reviewed study, and its 5% is not what the headline suggests

Marulanda, Lines, Kassa, Smithwick & Sullivan · Univ. of Kansas, Simplar Foundation, UNC Charlotte, Arizona State · published 2025-12-30 [13][14]

The only peer-reviewed measurement of an AI takeoff tool we could find compares Togal.AI with On-Screen Takeoff on architectural floor plans: a two-story fire station and a multistory hotel (two floor plans each), plus a control case and some initial tests on scanned sets, all performed by a single first-time user. It reports roughly 70% time savings (71% across the two main cases) and that “accuracy remained within a 5% margin compared to On-Screen Takeoff, with minimal discrepancies in smaller quantities.”

Read the method before you quote the number. The 5% is the percentage difference between Togal's values after manual adjustment and the manual On-Screen Takeoff values: the user corrected every quantity whose gap was 5% or more, then compared. The baseline is another tool's takeoff, not a surveyed ground truth, and the figure is post-correction agreement, not the AI's raw hit rate. The authors say so themselves: the single user may influence the results, running the tools in sequence may have favoured one over the other, lower-quality scans reduced accuracy, and relying solely on the AI-automated results “is not advisable.”

READ

The most careful measurement that exists, and it measures human-plus-AI agreement with a manual tool on clean drawings, then recommends review anyway.

TRADE PRESS · 1 ARTICLE

One six-tool “test”, with a scorecard of stars, not numbers

Edwards, Robotics & Automation News · 2026-02-19 · InEight, Togal.AI, Procore Estimating, STACK, Kreo, Beam AI [15]

The article most often cited as an independent benchmark says it fed six tools the same pack (more than 200 sheets, multi-discipline specs, a small Revit model, a stack of addenda) against a ground truth built by “senior quantity surveyors”, and scored them on a weighted card with accuracy at 40%. It reports InEight “1.8 percent” total error, Procore “within four percent”, STACK “within three percent” after assemblies were built for it, Kreo taking the testers “95 percent of the way”, and Beam with the “lowest miss rate after InEight” (no number).

For Togal it reports a twelve-minute takeoff, and for accuracy it relays the vendor's own figure, “the vision engine reports 97 percent”, yet its scorecard graphic marks Togal the accuracy category leader. The article names no testing organisation and no testers, and publishes no plan set and no per-tool error table; the scorecard itself is a grid of stars and shading whose only numbers are three callouts (InEight's 1.8%, Togal's twelve minutes, Beam's “lowest miss rate”).

READ

A narrative review with a decorative scorecard. Nothing in it can be checked or repeated, and its most-quoted Togal figure is the vendor's.

FIELD REPORTS · 2

The field numbers come with a drawing type attached, and swing 25 points inside one building

Jeppsen, Struvia · 2026-05-13 (secondhand, GC unnamed) · Drawer AI demo-webinar recording · May 2026 (vendor-measured) [16][7]

A May 2026 review quotes a general contractor on a Phoenix mixed-use job: by the GC's own estimate the tool ran at about 85% on the residential floors and about 60% on the retail podium level, and the podium correction “ate back half the time I saved upstairs.” The same review aggregates user reports of 85 to 95% on area takeoffs for clean multifamily, retail and office-TI PDFs, and worse on scans and MEP-heavy sets. Weigh the provenance: the GC is unnamed, the quote is secondhand, and the review sits on the blog of Struvia, itself an AI estimating product.

The other field number is a vendor's, given on a recorded sales call rather than a homepage: Drawer AI's distribution over 250 projects (above), with the presenter adding that at around 65% the QA effort eats most of the value of the automation, that the product is not tuned for single-family work, and that scanned drawings do not work at all. It is the only vendor statement on this page that quantifies where the tool fails.

READ

The only numbers with a drawing type attached. They vary by 25 points between floors of one building, which is the finding.

03 · Why they disagree

Five reasons a single percentage cannot be compared.

  1. Accuracy of what?

    Togal's figure is qualified “on floor plans”: spaces and areas. Drawer's is device counts. InEight's 1.8% is an aggregate quantity miss across a whole project. The Kansas study found its largest differences in small counts, not areas. A miscounted receptacle and a misdrawn room boundary are different failures with different costs, and a single percentage cannot carry both.

  2. Measured against what, by whom?

    A surveyed ground truth (the trade-press test, unpublished), another tool's manual takeoff (the Kansas study), the vendor's own databank (Drawer), or an estimator's recollection of the correction pass (the field reports). Four baselines, four different numbers for the same tool.

  3. On which drawings?

    The same tool on the same project: about 85% on the residential floors, about 60% on the retail podium. Clean vector multifamily sheets are the best case every vendor demos; scanned sets, dense MEP overlays and old as-builts are the sheets your bids actually contain. Drawer states plainly that scans do not run; Houzz limits AutoMate to simplified plans. Most vendors do not say.

  4. Before or after the human?

    The Kansas 5% is post-correction. Beam's 99%+ is post-human-QA, delivered in 24 to 72 hours. Togal's 98% does not say. Whether a number describes the machine's first pass or the reviewed deliverable changes what it means for your evening.

  5. Per sheet, or per bid?

    Two percent sounds like rounding. On the 760-device medical-office sheet in Drawer's own demo, it is about 15 devices, each one installed for free if it reaches the bid unreviewed. The number that matters is not the average; it is how many of the misses the tool pointed at.

04 · Before you trust a number

Six questions to ask any vendor.

Quote them. They are the same six we would want asked of us.

  1. “Ninety-eight percent of what: areas, counts, lengths, or dollars?”

    If the answer is not a unit, it is not a measurement.

  2. “Measured against what baseline, by whom, on how many sheets?”

    Ask for the plan set. A vendor that cannot name one has an impression, not a benchmark.

  3. “Is that before or after a human corrected it?”

    Post-review figures describe a service; first-pass figures describe the software. Both are fine, as long as you know which you are buying.

  4. “Show me the distribution, not the average.”

    Drawer's 15 / 60 / 90 split is more useful than any headline. A tool that is perfect on half your sets and useless on a quarter has a high average and a bad month.

  5. “How does it tell me which detections it is unsure about?”

    A first pass you must re-check in full saves nothing. A first pass that flags its own doubt lets you verify the flagged minority. This is the design question that separates tools.

  6. “Run it on my worst scan, before I pay.”

    One clean vector plan, your worst scan, and your densest electrical or MEP sheet. Count them by hand first; tally misses and false positives separately; time the correction pass. That is the only accuracy figure that applies to your shop.

05 · Where Bildrix stands

We publish no percentage. We show the doubt instead.

Bildrix does not publish an accuracy figure, and this page is the reason: no single number survives contact with a different office’s drawings. Rather than pick a flattering plan set and a flattering definition, the product is built around the failure mode every tool above has, so that when the first pass is wrong, you find out from the tool, not from the bid.

Every detection carries a confidence level. Low-confidence detections are flagged and queued for review, so you verify the flagged minority instead of re-counting the sheet. Every count, boundary and wire route is editable in the canvas, and what the AI misses you count or draw in the same takeoff. Messy scans and sparse sheets are labeled as such rather than reported as confident zeros. And the first takeoff is free on your own plan set, with a review call, so the measurement that matters is the one you make on your drawings.

Book your free takeoff reviewWith your review call · PDF plan sets
06 · Questions

Asked and answered.

How accurate is AI takeoff software?

No one outside the vendors has measured it in a way you could repeat. Advertised figures run from 94.5% to 99%+ with no dataset or definition behind them. The one peer-reviewed study found a single user's corrected first pass within 5% of a manual On-Screen Takeoff count on clean floor plans. Field reports range from about 60% on a retail podium level to the mid-90s on clean multifamily sheets. Accuracy is a property of the tool and the drawing together. Ask what a tool does when it is unsure, not what it averages.

Has anyone independently tested AI takeoff tools?

One peer-reviewed study (University of Kansas with Simplar Foundation, UNC Charlotte and Arizona State, published December 2025) compared Togal.AI with On-Screen Takeoff on four small projects with a single first-time user. One trade-press article (Robotics & Automation News, February 2026) describes a six-tool test but names no testers and publishes no data, and its Togal figure is the vendor's own. We found no public benchmark dataset for AI takeoff accuracy.

What does “within 5%” in the University of Kansas study mean?

It is the percentage difference between Togal.AI's quantities after the user manually corrected them and the quantities the same user measured by hand in On-Screen Takeoff. Corrections were made wherever the initial gap was 5% or more. It measures post-correction agreement between two tools, not the AI's raw hit rate against a surveyed ground truth, and the authors conclude that relying solely on the AI output is not advisable.

How should I test an AI takeoff tool on my own plans?

Pick three sheets you already know: one clean vector floor plan, your worst scan, and a dense electrical or MEP sheet. Count them by hand first. Run the tool, then tally misses and false positives separately and note whether it flagged its errors or presented them as certain. Time the correction pass. The figure that matters is corrected time against your manual time, on your drawings, not a percentage from anyone's homepage.

Does Bildrix publish an accuracy percentage?

No. There is no honest single number for a tool that runs on every office's drawings, and this page is the evidence. Bildrix shows its doubt instead: each detection carries a confidence level, low-confidence detections are flagged and queued for review, and every count, boundary and wire route is editable in the canvas. The first takeoff is free on your own plan set, with a review call, so you can measure it where it counts.

07 · Sources

Sources, in the order they are cited.

Every vendor page was fetched and read on 2026-08-19; quotations are verbatim from that read. Vendor pages change. If a figure above no longer matches its source, or we have misread one, write to support@bildrix.com and we will correct the page and date the change.

  1. [1]Togal.AI. Homepage: hero “up to 98%”; benefit block headed “Increase Accuracy” (figure: 98%, “on floor plans”); “5x faster”; “Takeoff in Minutes. Not Days.” https://www.togal.ai/
  2. [2]Togal.AI. Pricing page FAQ: “the fastest, most accurate estimating solution on the market” https://www.togal.ai/pricing
  3. [3]Kreo. “AI agentic workflow for takeoff and estimating” page: “Expert quality of 94.5-98.5%”; “up to 10,000 drawings per minute” https://www.kreo.net/solutions/ai-agentic-workflow-for-takeoff-and-estimating
  4. [4]Kreo. “AI construction takeoff software” product page: no accuracy figure on the page https://www.kreo.net/solutions/ai-construction-takeoff-sofware
  5. [5]Beam AI. Homepage: “99%+ accurate, QA-checked takeoffs and estimates in 24-72 hours”; webinar blurb “±1% of in-house accuracy” https://www.ibeam.ai/
  6. [6]Beam AI. “Beam AI vs Kreo” page: “100% automated takeoffs” and, in the same comparison table, “Every takeoff is reviewed by a human QA team” https://www.ibeam.ai/compare/vs-kreo
  7. [7]Drawer AI. Resources page: “Drawer AI Demo Webinar Recording” (posted May 2026); accuracy distribution stated at roughly 06:11, trust remarks at roughly 08:05 https://drawer.ai/resources
  8. [8]Bobyard. Homepage FAQ: takeoff time cut “50-90%, depending on drawing quality, trade, and the estimator review process” https://www.bobyard.com/
  9. [9]STACK. Floor Plan AI product page: no accuracy figure on the page https://www.stackct.com/floor-plan-ai/
  10. [10]On Center Software (ConstructConnect). Takeoff Boost page: “Highly Accurate AI… trained on millions of construction projects”; “up to 40% faster without compromising accuracy” https://www.oncenter.com/products/takeoff-boost/
  11. [11]PlanSwift (ConstructConnect). Homepage FAQ, “How accurate are the AI tools in Takeoff Boost?”, answered with “a head start, not a finished bid” https://www.planswift.com/
  12. [12]Houzz Pro. Help Center, “AutoMate AI for Takeoffs”: runs on “simplified floor plans and elevations only at this time” https://pro.houzz.com/pro-help/r/automate-takeoffs
  13. [13]Marulanda S., V. V.; Lines, B.; Kassa, R.; Smithwick, J.; Sullivan, K. “Togal.AI vs On-Screen Takeoff: A Comparative Analysis of Time Efficiency and Accuracy.” Journal of Research and Practice in Projects and Organizations (Simplar Foundation), published 2025-12-30. Full PDF. https://simplarfoundation.org/wp-content/uploads/2025/12/Togal.AI-vs-On-Screen-Takeoff.pdf
  14. [14]Togal.AI. Case-study page summarising [13] (“Peer Reviewed: Togal.AI vs On-Screen Takeoff”) https://www.togal.ai/case-study/peer-reviewed-study-togal-ai-vs-on-screen-takeoff
  15. [15]Edwards, D. “6 AI Construction Estimating Software Tested on Complex Project Accuracy.” Robotics & Automation News, 2026-02-19. https://roboticsandautomationnews.com/2026/02/19/6-ai-construction-estimating-software-tested-on-complex-project-accuracy/98967/
  16. [16]Jeppsen, B. “Togal AI Review: Is It Worth It for GC Estimators?” Struvia blog (Struvia sells AI estimating to GCs), 2026-05-13, updated 2026-07-16. https://struvia.co/blog/togal-ai-review-2026

Method: a claim appears in Table 1 only if we found it on the vendor’s own site; the measured section includes every independent or semi-independent measurement of a named AI takeoff tool we could locate as of 2026-08-19. Bildrix competes with the vendors in Table 1; read the page with that in mind. Published 2026-08-19; last revised 2026-08-19.