True Data · Methodology

How True Data works

True Data provides on-demand benchmarks backed by a corpus of normalized program, design, cost, schedule, risk, and decision data from relevant commercial construction projects. It lets owners and the project teams who serve them compare a project with industry norms to facilitate a shared understanding of key outcome drivers.

This page describes the methodology for constructing the benchmark values and how that data is handled and kept private in anonymized market-data benchmarks.

Where we are today: we're bringing projects online in a rich, standardized format to enable reliable benchmarking. The additional privacy protections described below are designed to be in place before general market data is available for members to use.
00 · What it does

Predict ranges and compare with relevant cohorts

Every benchmark is relative to one or more cohorts of comparable projects. Depending on what you start with, a benchmark report will:

  • Predict ranges. True Data provides distributions for estimated ranges implied by what you currently know about the project, combining statistical inference from cohort distributions with “stretch-to-fit” scaling of comparable projects' design breakdowns to your known dimensions.
  • Compare project values to cohort-derived ranges. When you have a value for your project, e.g. a skin-to-floor ratio or an estimated cost per square foot, we show where it falls in the distribution of comparable projects and how the cohort set compares with your project across both quantitative and qualitative criteria.

Cost is normalized to the UniFormat Level 2 elemental level and, where needed, escalated for time and location to make projects comparable. We stop there for anonymous market data we do not show detail below Level 2.

01 · What we benchmark

Program, design, schedule, cost, risk, and decisions

A benchmark spans the dimensions that capture project outcomes, each compared against the cohort and most often shown as a box-and-whisker or violin distribution rather than a single number.

  • Program. Scale and mix of functional areas of the project.
  • Design efficiency. A set of commonly referenced ratios that signal how efficiently the design expresses the program: skin-to-floor ratio, grossing factor, window-to-wall. What is relevant and compared depends on the use types in your project.
  • Cost. Normalized to company standards and UniFormat Level 2 (for market data), composed to match the program of your project, and escalated for both time and location.
  • Schedule. Major schedule milestones and the phase durations between them.
  • Risk and Contingency. Register entries scored by likelihood and impact to cost and schedule, with likely-relevant risks surfaced based on the project profile and contingency distributions by phase.
  • Decisions. Critical decisions by project phase.
  • Sustainability Embodied carbon intensity (up-front and whole-life), energy and water use intensity.

Alongside these, we capture the program requirements that drive cost and schedule — things like a LEED certification target or finish quality level. These condition the cohort and help explain why a project's numbers differ from a naive comparable.

Benchmark report

Design efficiency
Skin-to-floorIn range
Window-to-wallAbove range
Floor-area ratioIn range
Cost$410 median $/GSF
ABCDG
Schedule · months
01224mo
Cohort composition28 projects
By type
By region
By delivery
Blue = cohort · Plum = yours
02 · Where the data comes from

Benchmarks are only as good as what goes into them.

The corpus is built from the documents project teams already produce at their design and procurement milestones. Every contribution is conditioned and qualified before it can back a benchmark.

Data points come from the primary artifacts a project team generates at a GMP or other design-phase milestone, the same records the team relies on and uses to communicate with the owner as the project progresses:

  • Program & requirements documents. The owner's program, basis of design, and requirements — unit and bed counts, department areas, finish quality, and certification targets like LEED — that explain why a project's numbers land where they do.
  • Estimates Cost of Construction estimates which we normalize to UniFormat Level 2; owner cost models including land acquisition, financing, soft costs, and FF&E.
  • Schedules Milestones and the durations between them extracted from the project schedule.
  • Models & Drawings 3D models, architectural and engineering drawings, from which the design-efficiency ratios and program quantities that other sources don't capture directly are derived.
  • Specifications The written specifications that accompany the drawings detailing the materials, systems, and quality standards. These qualify the finish and performance level behind a project's cost and design figures.
03 · Data layers

Your data, peer data, and market data

A benchmark can draw on up to three sources of data. They have the same underlying shape but differ in how much of each underlying project you can see.

  • Your data.Full detail Projects from your own organization. For these project you see everything and can inspect both the processed records and the primary sources. You may have additional cost codes or other data associated with your organization's own projects that you use in benchmark reports but that remains private to other members of your organization.
  • Peer data.Individual points If you belong to a peer group whose members have agreed to share with each other, you see the group's projects as individual data points — but only the fields captured in the standardized contribution format, and nothing beyond that. Counts are exact, since a group you're part of isn't anonymized to you; the distributions still pass through the same privacy protections as market data.
  • Market data.Aggregate only Anonymized, aggregated data from every contributing organization. It appears only as aggregate distributions across cohorts of comparable projects, never as individual, identifiable projects. It is the largest set, and the privacy protections below govern it.

The more of these a benchmark can draw on, the more comparable projects stand behind it — but only your own data is ever shown in full.

04 · How it's prepared

Conditioned and Qualified

Benchmarks are only as good as what goes into them. The corpus is built from the documents project teams already produce at their design and procurement milestones. Every contribution is conditioned and qualified before it can back a benchmark.

A processing pipeline conditions data from primary sources into a consistent, comparable shape, mapping costs to a common elemental taxonomy, aligning schedule milestones, and tagging the program drivers that define a project's cohort. This normalization is reviewed by people with deep domain expertise, and the conditioned project data is available through an API for use in other applications. This conditioning is offered as a service to address the (often substantial) burden of data conditioning that makes it difficult for organizations to use their own historical data confidently.

Qualification is the gate that decides whether a conditioned project datapoint is sound enough to be included in market benchmarks: we check that its milestone, estimate, and program are internally consistent and complete before it's allowed to influence any distribution. Datapoints that can't be reconciled are held back from anonymized market cohorts rather than allowed to distort a benchmark. This is applied by default to own-data cohorts as well, but you can still access the partial datapoint and choose to use these projects where appropriate

Want the specifics? See the data format every contribution is normalized to.

04 · Dynamic cohorts

Cohorts are fuzzy, interactive, and comparable side by side

A comparable project doesn't have to match yours on every attribute. We seed one or more default cohorts from a project's known attributes, then let you adjust them in the report itself to match what's important for your project.

  • Fuzzy membership. Cohorts are seeded from relevant attributes like type, sector, region, delivery method, finish level, and floor-area range that correlate with benchmarked values. A cohort is not an exact match on every parameter, and you can inspect cohort composition for each attribute.
  • Interactive slicing. Drill into an attribute donut chart (e.g. sector or wage basis) or filter a shown distribution (e.g. total GSF) to refine a cohort or build a new one. Slice however you like for your own data, or use defined privacy-controlled cohorts with market data.
  • Side-by-side comparison. Compare multiple cohorts at once — “union vs. open shop,” or “this metro vs. national” — to see how a single attribute moves the numbers.

A benchmark is only as useful as your ability to understand what kind of projects back it, so cohort composition including attribute distributions, and approximate size are visible.

05 · Privacy

Anonymized market data protects privacy

In market data benchmarks, a contributing project is never shown to other users directly. It is surfaced only in aggregate distributions across cohorts of similar projects, never as identifiable individual values. We're building the privacy layer on the techniques of differential privacy — the approach used by the U.S. Census Bureau and Apple — which adds calibrated noise and caps how much any single project can affect what's shown. That is what prevents reconstructing an individual project's underlying information from any benchmark view or combination of views while allowing members to draw on a broad industry data set to power their analysis.