Seven layers, a trust state machine, and 100 TB of combined ground truth — 35 TB of it Android — behind every score we emit. This is the deep end. The products it powers live on the main page.
Evidence flows down from devices to decisions; intelligence flows back up, because every evaluation refines the baselines. The order below is the order the data travels.
Harvests thousands of high-volume signal types at three timescales — keyboard dwell & flight dynamics, file create/delete/rename rates, extension entropy, ransomware-like burst patterns, process creation trees, DNS resolution behavior, domain entropy & novelty, and session-level browsing abstractions — with integrity context (missingness ratios, agent health) attached to every batch.
120s / 900s / 3600s multi-timescale windows2,100+ signals per agent, six signal familiesmetadata-only network side — privacy by designConverts noisy real-world activity into a mathematical representation of behavior: aggregation, extraction of behavioral, statistical and temporal patterns (cadence, entropy, sequence, distributional moments, trends, bursts, periodicity), then missing-resilient normalization so degraded telemetry never poisons a score.
1,024+ feature vector per entity per window+100–250 integrity / meta signalsMIDAS-based missing-aware normalizationLearns what normal looks like — per entity, per session, per host, per tenant, per fleet. Exponentially-weighted baselines adapt as the fleet drifts; a neural autoencoder learns each entity's behavioral manifold; embedding space is monitored so concept drift is caught before it becomes blindness.
z₁ = αx₁ + (1−α)z₀ EWMA adaptive baselineEnc(X)=Z · Dec(Z)=X̂ autoencoder manifoldCSV core registry versioned by registry hashQuantifies how far current behavior sits from the learned signature, five ways at once: statistical residual deviation, change-point detection for abrupt shifts and gradual drift, reconstruction error against the learned manifold, KL divergence between embeddings, and distance from the peer/fleet baseline — with EVT/GPD adaptive thresholds for genuinely novel behavior.
Z₁ = (x₁ − μ)/σ residual deviationPage-Hinkley · CUSUM change-point & driftE₁ = ‖X − X̂‖² · Dₖₗ(P‖Q) manifold & embedding driftSynthesizes divergence signals, graph data (inter-sensor linkages, similarity and causality, peer baselines) and contextual evidence into one coherent risk view — probabilistic fusion weighs each signal by significance and influence, and temporal-graph relevance weighs graph changes by recent drift history.
dynamic Bayes probabilistic signal fusionoutlier rank similarity + meta-signal ensembleJaccard · FCD fleet clustering & campaign likelihoodOne score in [0,1], recomputed continuously, with a priority level, a transparent decomposition of exactly which signals moved it, and suppression/policy gates so 150 suppression rules keep noise from becoming alert fatigue. Thresholds trigger alerting, blocking and contextually-aware response — detect, monitor, explain, act.
S₁ ∈ [0,1] descriptive confidence attachedLow / Med / High / Critical priority levelsexplainable transparent model decompositionTurns scores into safe action: suppression rules and policy gates keep noise from becoming alert fatigue, trust-score thresholds trigger alerting, blocking and contextually-aware response — and every decision feeds back to retrain baselines, closing the loop between detection and defense.
150 suppression rules noise kept below alert fatiguealert · block · respond fusion-based action on thresholdsfeedback loop every decision retrains the baselinesRather than a binary verdict, each entity holds a modelled trust state that decays as anomalies accumulate and recovers as behavior returns to baseline — every transition attributable to understandable contributors.
TrustGate isn’t a scanner with a UI — it’s a research program shipped as a product. These are the principles every layer is built to honor.
Verdicts anchor in hardware attestation verified off the device and network metadata observed from outside — signals a resident implant cannot forge, no matter how deep it sits.
Deterministic rules beside deep neural detectors, trained and continuously retrained on 1.5M+ labeled samples with a matched goodware set — driving false positives and false negatives down together, not one at the other’s expense.
A 15-dimension fusion engine produces one calibrated score, decomposed into contributors and mapped to MITRE ATT&CK Mobile — so every verdict arrives with its reasoning attached.
Confidence stated, limits declared, blind spots documented. Decision support, not oracle theatre — the property that governments, journalists and executive-protection teams actually buy.
The pipeline that turns 100 TB of combined ground truth — 35 TB of it Android — into deployed intelligence. Abstracted by design; specifics live under NDA.
Any competent engineer can rebuild collection. What cannot be rebuilt is a decade of labeled ground truth — and a scoring stack trained against it.
Intelligence is only as trustworthy as its sources. Every signal in our stack traces to one of six documented origins — collected read-only, consent-based, and metadata-only on the network side.
1.5M+ Android malware samples gathered continuously since 2016 — ~35 TB of APKs, each carrying multi-engine detection metadata, family evolution history and signing-certificate lineage.
A matched known-good corpus built from CIRCL's public hashlookup service — the negative signal that proves a binary clean, offline, and kills false positives instantly.
Continuously refreshed open-source IOC feeds — hashes, domains, IPs — merged into the intelligence database, so today's campaigns are matched against today's indicators.
Google's current attestation roots, tracked through the 2026 root rotation, so both established and newest devices verify. Evidence signed by secure hardware — not vouched for by software.
Every TrustGate assessment and every fleet-agent window writes back into the shared corpus — baselines, fuzzy indices, verdicts — under explicit consent, read-only on the device.
A label-capture loop records analyst verdicts, retroactive corrections and survival labels on every reviewed finding — the training signal that keeps deterministic rules and ML models honest.
Introduces continuous behavioural trust modelling — integrating baseline statistics, change-point detection, neural autoencoders and embedding-drift analysis into a single continuous trust score — and shows earlier detection of credential abuse, slow insider threats and AI-assisted attacks than event-based approaches.
A formal, privacy-aware, cross-platform semantic model for trust-relevant signals: a clean signal universe, entity model, canonical naming, causal relationships and trust-state definitions with decay and recovery contributors — the foundation for explainable trust scoring.
Research access on request — write to the lab →
Compromise assessments, design partnerships for the Behavior Trust Layer, corpus and API licensing, or an investor briefing — tell us which, and we'll come prepared.