Skip to content

Validation / Methodology

Methodology, evidence, and the road to proof

The full record behind the read, written by the side trying to falsify the product. It scores every dimension of validity honestly, reports the failures alongside the passes, and lists the pre-registered program and owner gates that would earn the word impact.

What you can rely on today

Five things are true right now, before a single further study runs. Each is stated at the strength the evidence supports, and every number below traces to a named source.

A reliable, reproducible engine

One deterministic scoring core, ported and checked to the integer (78 of 78 parity) and stable under small perturbation. The same footage returns the same read, every time.

We measure what we can see, and say so

The engine reads the room at the group level: where attention points, motion, expression geometry, and how in sync the room is. Aggregate only. No individual is scored, and no emotion is inferred.

Private by construction

Footage is processed and deleted as soon as the analysis finishes. No faces are retained, only the derived aggregate numbers, so there is no individual record because none is ever created.

In step with the science and the regulators

The signals are grounded in the facial-action and behavioural-measurement literature, and the engine infers no emotion, consistent with the regulator position (EU AI Act Article 5(1)(f)) that emotion recognition lacks a scientific basis. A runtime gate fails closed on emotion-framed output.

Proven in the open, not asserted

Every claim here is stated at the strength the evidence supports, with named sources. The harder claim, that a read predicts a real outcome, is held behind a pre-registered program rather than asserted. We hold the word until it is earned.

What we are not claiming yet

That a Proof of Impact score predicts a real-world outcome you would call impact. Today that is a measured appearance correlate and a hypothesis under test, not a proven result. The full, unsoftened validity scorecard and the pre-registered program that will test it are below, in the open.

What the word Proof can and cannot mean here

The product is named Proof of Impact. Here is the honest split between what is proven, what is only measured as an appearance correlate, and what is still a hypothesis.

Proven

That the software is reliable and reproducible: one deterministic, faithfully-ported scoring core (78 of 78 parity), stable to small perturbation. This is about code, not about what the scores mean.

Measured

Behavioural appearance correlates: where faces point, motion, expression geometry, and room synchrony. These are aggregate, group-level reads. They are measurements of appearance, not of attention, emotion, or impact.

Hypothesized

That these correlates track, or predict, a real outcome anyone would call impact. No study links a score to an outcome yet, so this is a hypothesis, not a result. Until it is tested, impact is not proven.

Validity scorecard

Every dimension of validity, scored honestly. A Not established is never softened into a Partial. This is the spine of the dossier and of the go or no-go on the landing-page link.

DimensionStatusWhy
Construct validityNot establishedThe engine measures appearance correlates, not the latent construct impact or engagement. The one external-criterion test of its expression substrate (FACS AU4/AU6/AU12 versus OpenFace) fails to converge (CCC below 0.6 on all 9 clips), so any expression-derived read is a behavioural correlate, not a calibrated construct measurement.
Content validityPartialThe aggregate attention, energy, and synchrony signals are grounded in cited behavioural-science literature and the product scopes its claims to correlates, which supports partial content validity. It is partial and not established because no independent content-expert panel has judged coverage, and the expression component inherits the failed FACS convergence.
Criterion and predictive validityNot establishedNo study links any score to a real outcome. The Altitude venue sessions carry no outcome measure, so this cannot move on that footage. A commissioned predictive protocol is published; until an outcome study runs and passes its pre-registered gate, this stays Not established. This is the dimension that would license the word impact, and it is not earned.
Ecological and external validityNot establishedAlmost all evidence is benchmark, studio, or semi-synthetic. On the single real-venue test the skin-tone proxy demonstrably collapsed. Evidence that does not transfer to the deployment domain is not validation of the deployment domain. The venue program is pre-registered but blocked on consent, compute, and a working GPU pipeline, so no venue study result exists yet.
Reliability (measurement stability)PartialInternal reliability is shown: internal consistency (sub-score correlations with expected signs) and stability to 5 percent landmark noise, plus deterministic, reproducible output. But split-half, camera-concordance, and test-retest reliability on venue footage are pre-registered and unrun, so measurement reliability in the deployment domain is not established. Reliability is not validity.
Software reliability and reproducibilityEstablishedThe shipped TypeScript core matches the canonical Python engine on 78 of 78 cells exact (the 5 core plus single-person indices), output is deterministic (0 variance over repeated runs), and the surfaces conform to 0 delta on the FACE-ONLY core at a shared frame rate. Two scope limits: parity does NOT cover the multi-person aggregate path (deriveMultiPersonScores) that produces the shipped room read, and body-dependent scores still diverge up to 10 points across surfaces because body runs only on web (D8, open). This is Established as software reliability and reproducibility ONLY. It is explicitly not construct, criterion, or ecological validity: two programs agreeing on a number is not the number being right.
Measurement invariance and fairnessNot establishedMST-E studio portraits show a real darker-skin detection deficit around the 20 px floor (studio, and the reported intervals overstate precision until a person-level recompute runs). On real venue footage the tone proxy collapsed entirely, so no venue fairness gap is measurable today. The mitigation is studio-validated only. Venue fairness is Not established and is a pre-registered study.
Adversarial resilienceNot establishedThe engine carries abstain and confidence logic by design, but it has not been adversarially tested against the pre-registered negative controls (empty room, temporally shuffled frames, a single looped still). Until those controls run and the engine is shown to abstain rather than emit confident reads on them, adversarial resilience is Not established.
IndependenceNot establishedEvery result here is vendor-run. There is no third-party or academic replication of any figure. Independence is Not established and the page claims no external validation until it exists.

Claim by claim

Each card states the claim at the strength the evidence actually supports, with its method, effect size, honest uncertainty, the threats a hostile expert would raise, and the named sources every number traces to.

The shipped single-person scoring math is a faithful, reproducible port of the canonical engine.

Established

Dimension: Software reliability and reproducibility

Evidence
The boost-app TypeScript scoring core matches the canonical Python engine on 78 of 78 cells, exact to the integer, across every canonical scenario and context (the 5 core plus 8 single-person indices). Output is deterministic: 0 variance over 10 repeated runs of the full pipeline. The three surfaces conform to 0 delta on the FACE-ONLY core at a shared frame rate.
Method
A parity shim runs the shipped core and the canonical Python engine on shared single-person fixtures and compares all 13 score fields cell by cell; determinism and conformance are measured by repeated and cross-surface runs on fixed input.
Effect size and uncertainty
78/78 exact match on the 5 core plus single-person indices; 0.0 variance; 0.0 FACE-ONLY cross-surface delta at a shared frame rate.
Threats to validity
This is software reliability and reproducibility, not validity. Two scope limits carry forward: (a) parity covers the 5 core and single-person indices only. The multi-person AGGREGATE scoring path (deriveMultiPersonScores), which is the shipped Proof of Impact room read, is NOT covered by the 78/78, and its parity is a separate, unrun exercise. (b) Conformance is 0 delta on the FACE-ONLY core; body-dependent scores diverge up to 10 points (Stress most) across surfaces because body extraction runs only on web (D8, open). Two programs agreeing on a number is not the number being right.

Sources: docs/validation/engine-parity.md, docs/engine/02-conformance.md, docs/validation/03-local-validation-results.md

The engine is internally consistent and stable to small perturbation, but its venue reliability is unmeasured.

Partial

Dimension: Reliability (measurement stability)

Evidence
Sub-score correlations carry the expected signs and strength (Composure to Stress r = -0.92, Momentum to Engagement r = +0.83). Under 5 percent landmark noise, 16 of 16 scores stay within 15 points, most moving 0 to 2. Split-half, camera-concordance, and test-retest reliability on real venue footage are pre-registered but not yet run.
Method
Internal consistency and noise sensitivity are computed on synthetic scenarios in the local validation suite; venue reliability studies are specified in the pre-registration and are pending.
Effect size and uncertainty
r = -0.92 and +0.83 on the two headline pairs; 16/16 scores stable within 15 points under 5 percent noise. Venue split-half ICC: not yet measured.
Threats to validity
Internal consistency and noise stability are reliability, not validity, and are measured on synthetic inputs. Deployment-domain reliability (split-half on real footage, camera concordance) is the open piece.

Sources: docs/validation/03-local-validation-results.md, docs/validation/altitude-preregistration.md

On a crowd detection benchmark, recall falls off below about 20 px and tiling roughly doubles crowd recall. This bounds detection on a benchmark, not on the venue.

Not established

Dimension: Ecological and external validity

Evidence
On WIDER FACE, tiled recall is 0.64 at 15 to 20 px, 0.78 at 20 to 30 px, and 0.88 at 60 px and up; tiling lifts dense-crowd recall from 0.19 to 0.43 at 101 or more faces per image. The room count is therefore a detected lower bound that undercounts dense rooms by roughly half.
Method
Recall at IoU 0.5 on a stratified 400-image sample (21,201 faces) of the WIDER FACE validation set, single-pass versus tiled, by face size and image density.
Effect size and uncertainty
Tiled recall 0.64 (15 to 20 px) to 0.88 (60 px and up); dense-crowd recall 0.19 to 0.43 with tiling. Faces cluster within images, so intervals should be image-clustered.
Threats to validity
WIDER FACE is a public benchmark, not the deployment venue, and has no demographic labels, so it supports no ecological-deployment or fairness claim. The size and density envelope is a benchmark measurement only.

Sources: docs/engine/crowd-scale-result.md

On real venue footage the skin-tone proxy collapsed, so a venue fairness gap cannot currently be read. This is a published failure.

Not established

Dimension: Measurement invariance and fairness

Evidence
On the Altitude footage every harvested face was classified into a single tone band (proxy_valid: false); median ITA around -71 to -73, far past any real skin value, because indoor GoPro exposure crushed luminance. Even a normalization pass did not restore a trustworthy tone axis.
Method
Pseudo ground truth built from the footage: detect large confident faces, tag tone by ITA proxy on the crop, then re-detect at rendered crowd sizes. The tone axis was checked for validity, not assumed.
Effect size and uncertainty
Populated tone bands: 1 (degenerate). Median ITA around -72, outside the plausible skin range. No venue per-tone gap is computable as-is.
Threats to validity
This is the honest negative result and the reason the venue fairness program exists. Studio MST-E fairness does not transfer here. Closing this is pre-registered Study 8, gated on consent and compute.

Sources: docs/engine/altitude-footage-fairness-result.md, docs/validation/altitude-preregistration.md

On studio portraits with ground-truth Monk labels, the detector loses darker faces faster at the crowd-critical size. This is studio, not venue.

Not established

Dimension: Measurement invariance and fairness

Evidence
On MST-E, at the 20 px floor darker-skin recall is 0.52 against 0.74 for lighter faces (gap about 0.22); parity arrives by 30 px, and all tones collapse together by 15 px. A CLAHE plus upscaling mitigation lifts darker 20 px recall to about 0.70 and halves the gap.
Method
Detect each MST-E portrait, tag by ground-truth Monk band, render at crowd pixel sizes, re-detect; recall per band per size over the full 1543-image run.
Effect size and uncertainty
Darker recall 0.52 vs lighter 0.74 at 20 px (image-level; person-level interval pending). Effective sample about 19 people, not 1543 images.
Threats to validity
Pseudoreplication: the reported intervals treat 1543 images as independent when there are about 19 people; a person-level recompute is required and gated on re-acquiring the licensed data. And this is studio, not venue: on venue footage the proxy collapsed. Not established for the deployment domain.

Sources: docs/engine/mste-fairness-result.md

Cross-camera identity does not hold when cameras see different people; the multi-camera claim is not shipped.

Not established

Dimension: Reliability (measurement stability)

Evidence
On 9 people with a synthetic flipped second view, the disjoint-population false-merge rate is 2.8 percent at the deployed threshold and up to 26 percent at the low end, and the hardest same-demographic pair is unseparable at any usable threshold. Detection recall (0.933) clears its floor; the identity claim does not.
Method
Full enumeration of all 252 identity-disjoint splits of the 9 people, EER threshold on a calibration split, reported on held-out; plus a live end-to-end run on Modal.
Effect size and uncertainty
Disjoint false-merge 2.8 to 26 percent (does not clear a 1 percent bound). Effective sample 9 people; the 252 splits are correlated, not 252 samples.
Threats to validity
Pseudoreplication (9 people, 252 correlated splits) and a synthetic second view (a horizontal flip, not a real angle). The honest recommendation is to not ship the multi-camera identity claim; it waits for labeled multi-view data.

Sources: docs/engine/sface-validation-result.md

The expression substrate does not converge with an independent FACS criterion, so expression-derived reads are correlates, not calibrated measurements.

Not established

Dimension: Construct validity

Evidence
Against OpenFace, AU4 and AU6 concordance (CCC) stays below 0.6 on all 9 clips; AU12 tracks direction well (Pearson up to 0.96) but fails calibrated CCC. Zero of nine clips pass the priority-AU gate.
Method
Concordance correlation (CCC, gate 0.6) between the engine AU substrate and OpenFace on the convergence corpus; the corpus itself rates only 3 of 9 clips as usable material.
Effect size and uncertainty
CCC below 0.6 on all 9 clips for AU4 and AU6. The one external-criterion test, and it fails.
Threats to validity
The corpus is weak, which confounds the result, but even on the usable clips AU4 and AU6 do not converge. The scores remain behavioural correlates, which the product already states.

Sources: docs/validation/03-local-validation-results.md

The multi-hour capability claim is not yet substantiated on venue footage; no continuous two-hour session even exists in the corpus.

Not established

Dimension: Reliability (measurement stability)

Evidence
The read-only inventory shows 10 distinct capture sessions across 2 days, with the longest continuous single recording at 50.2 minutes and 0 clip groups safely concatenable. A two-hour endurance run would be an assembled input, not a continuous session, and is blocked on compute and a working chunked GPU pipeline.
Method
ffprobe metadata inventory plus creation_time clustering (no frame decoded); the multi-hour pipeline study is pre-registered as an endurance and cost test.
Effect size and uncertainty
Longest continuous view 50.2 minutes; 0 of 13 clip groups concatenable; multi-hour run not executed.
Threats to validity
Endurance is not a validity claim about the read. The corpus cannot supply a genuine continuous two-hour single-audience recording, so any multi-hour input is a stitched endurance harness, labeled as such.

Sources: docs/validation/altitude-corpus-manifest.json, docs/validation/altitude-preregistration.md

The shipped read is aggregate attention, energy, and synchrony, framed as behaviour and geometry correlates, with no emotion inference.

Partial

Dimension: Content validity

Evidence
The aggregate room report emits only attention on the speaker, room synchrony, a session-phase energy arc, and behavioural pillars, with no per-person affect. An EU AI Act Article 5(1)(f) runtime gate fails closed on emotion-framed output in workplace and education contexts.
Method
Allowlist projection in the aggregate report plus a runtime emotion gate with unit tests; the framing is grounded in cited behavioural literature.
Effect size and uncertainty
No numeric effect size; this is a scoping and compliance posture, grounded but not validated by an external content-expert panel.
Threats to validity
Content grounding is not construct or criterion validity. The framing reduces legal and overclaiming risk; it does not prove the correlates measure the underlying construct.

Sources: docs/legal/access-compliance.md, app/what-it-measures/page.tsx

No score has ever been linked to a real outcome, so the product cannot honestly use the word impact yet.

Not established

Dimension: Criterion and predictive validity

Evidence
No criterion or predictive study exists. The Altitude sessions carry no outcome measure (session ratings, evaluations, net promoter, recall, behaviour). A predictive protocol is commissioned and published; it runs only if outcome data is found for these sessions.
Method
Gap statement. The commissioned predictive protocol is specified in the results document.
Effect size and uncertainty
No effect size exists because no outcome variable has been linked.
Threats to validity
This is the load-bearing gap behind the product name. Until a pre-registered predictive study passes, criterion and predictive validity is Not established and the page says so.

Sources: docs/validation/altitude-results.md, docs/validation/altitude-preregistration.md

Every result here is vendor-run; there is no independent replication.

Not established

Dimension: Independence

Evidence
No third-party or academic group has replicated any figure on this page. The program is designed to be replicable (pre-registered thresholds, named harnesses, committed artifacts), but replication has not happened.
Method
Gap statement.
Effect size and uncertainty
Zero independent replications.
Threats to validity
Vendor-run validation is the weakest form of assurance. The page claims no external validation until an independent review exists.

Sources: docs/validation/validity-scorecard.json

The engine has not been shown to abstain on negative controls; that test is pre-registered and unrun.

Not established

Dimension: Adversarial resilience

Evidence
The pre-registration requires the engine to abstain or return null on an empty room, a temporally shuffled session, and a looped still frame. These have not been run on venue footage. If the engine emits a confident room-energy or attention read on any of them, that is disconfirming and will be published.
Method
Pre-registered Study 10; not yet executed (gated on consent and compute).
Effect size and uncertainty
No result yet; required behaviour and pass or fail thresholds are fixed in the pre-registration.
Threats to validity
Abstain logic existing in code is not the same as passing an adversarial test. Until the controls run, adversarial resilience is Not established.

Sources: docs/validation/altitude-preregistration.md

How a skeptic would attack this, and our answer

The strongest version of the case against this product, stated by us, with our honest answer. If any answer reads as a dodge, it is a bug in this page.

You call it Proof of Impact, but you never measured impact on anyone.

Correct, and we say so at the top. The engine measures behavioural appearance correlates (where faces point, motion, expression geometry, synchrony), not impact. Criterion and predictive validity is Not established, and no score has been linked to a real outcome. The name is reconciled in the proof-measured-hypothesized panel above.

Your evidence is benchmarks, studio portraits, and synthetic data, not the venue you deploy in.

Also correct, and it is why the venue program exists. On the single real-venue test the skin-tone proxy collapsed. Evidence that does not transfer to the deployment domain is not validation of the deployment domain, so ecological validity is Not established until the pre-registered venue studies run and pass.

Your ground truth is a human watching the same faces the engine watches, so agreement is circular.

We flag the circularity in the pre-registration. Coders judge appearance and the engine measures appearance, so convergence is evidence about a shared appearance construct, not proof of attention or impact. It is convergent evidence with a stated ceiling, never a criterion for impact.

Your samples are tiny and clustered: one event, one audience, a handful of people.

The effective sample is about 10 capture sessions inside ONE event with one recurring audience, the cross-camera work used about 9 people, and MST-E has about 19. Our reporting unit is the session or person, and we report the effective n in sessions and people, never in frames. The session-level cluster bootstrap and the multiplicity correction are the pre-registered method for the venue program, which has not run; the studies that HAVE run report image- or split-level intervals that we explicitly flag for a required, still-pending cluster-aware recompute (for example the MST-E person-level recompute).

Two code paths matching to the integer is not science.

Agreed. The 78 of 78 parity, determinism, and conformance results are reclassified as software reliability and reproducibility, not validity. Two programs agreeing on a number is not the number being right, and the build check fails if any real-world-validity claim rests only on parity.

Confounds: lighting, camera angle, seating depth, session content, and culture all move your signals.

They do. The sensitivity study varies the arbitrary knobs, the negative controls test whether the engine abstains on empty rooms and shuffled frames, and the known-groups check ties signals to real events. None of these has run yet, so any read remains confounded by lighting, angle, seating, content, and culture until they do.

No one independent has checked any of this.

True. Every figure is vendor-run. Independence is Not established, and we claim no external validation. The program is built to be replicable (pre-registered thresholds, named harnesses, committed artifacts) precisely so an independent group can check it.

Regulators say emotion inference from faces has no scientific basis.

We agree with that position and do not infer emotion. The EU AI Act Article 5(1)(f) and Recital 44 cite the lack of scientific basis for emotion recognition; a runtime gate fails closed on emotion-framed output in EU and UK workplace and education contexts. The read is aggregate attention, energy, and synchrony, a correlate, not an emotion.

The venue-footage program

Real audiences in a real venue: the domain the current evidence lacks. The program is pre-registered before any result, so the thresholds cannot be moved to fit the data. It has not run: it is blocked on the owner gates below.

10

distinct capture sessions, one event, one recurring audience

0/10

pre-registered studies passed

50.2m

longest continuous view (no continuous two-hour session exists)

Effective sample size is about 10 distinct capture sessions inside ONE event with one recurring audience, not 39 files and not the millions of frames. Clip-number grouping over-merges (simultaneous multi-camera captures), so no continuous two-hour session exists in clean form.

Read the detail:docs/validation/altitude-preregistration.mddocs/validation/altitude-results.mddocs/validation/altitude-corpus-manifest.json

Status of the validation gate

Software reliability is established today. The deployment-domain gates (ecological, criterion, and fairness validity) are pre-registered and still pending, so the stronger claim that a read predicts impact is not yet earned. This status is published in the open, never hidden.

  • Venue studies not passed (0/10 pass, rest pending or failed).
  • Consent basis for research analysis of the footage is not documented.
  • No deployment-domain validity dimension (ecological, criterion, or fairness) is Established.

This status is driven by the committed results artifact (docs/validation/validity-scorecard.json), not a hand-set switch. Artifact version v0.1.0, dated 2026-07-08.

Owner gates

These are not engineering choices. They are decisions and reviews only the owner and counsel can close, and they block the studies, the publication, and the word impact.

open

Consent basis for research analysis of the Altitude footage, plus retention and deletion schedule

Blocks: all venue studies

open

Counsel review and sign-off before publish

Blocks: publishing the dossier beyond an internal artifact

open

Independent third-party or academic review

Blocks: any claim of external validation

open

Whether outcome data exists for the Altitude sessions

Blocks: criterion and predictive validity

open

Compute budget for the multi-hour and full-corpus runs, and the go-live decision

Blocks: multi-hour and full-corpus studies

Expert appendix

Statistics posture. The reporting unit is the session or person, and the effective sample size is reported in sessions and people, never in frames or chunks. The session-level cluster bootstrap and the Benjamini-Hochberg multiplicity correction are the PRE-REGISTERED method for the venue program, which has not run. Intervals on studies that have already run (for example MST-E) are image or split level and are flagged for a required, still-pending cluster-aware recompute; they are not presented as person-level confidence.

Errata. The prior evidence was audited and corrected: software parity, conformance, and determinism were reclassified as software reliability, not validity; MST-E intervals were flagged for pseudoreplication (about 19 people, not 1543 images) with a person-level recompute marked as the required, still-pending next step; the 57-face MST-E pilot figures are superseded; and the multiplicity and effect-size framing was added across the tone and size comparisons. The errata notes live in each corrected document.

No fabricated authority. There is no certification, no peer review, and no independent audit of this work. The engine infers no emotion, consistent with the regulator position that emotion recognition lacks a scientific basis.

Full record. The pre-registration, the corpus manifest, the full results, the validity scorecard artifact, and each corrected source document are the authoritative record behind every figure on this page.