Dotsin Research Labs

Modelling the
human,
not the text.

We build the Large Behavioural Model (LBM) — the world's only foundational AI architecture that treats human behaviour as a first-class computational primitive.

Mission & Vision
Why we exist.
Language has LLMs. Vision has ViTs. The physical world has JEPA. Human behaviour has no foundation model. We're here to change that.
Mission
Build the world's only foundational AI that understands human behaviour — not through clicks and keystrokes, but through the causal, temporal, multi-dimensional structure of how humans actually function.
Vision
A world where every AI system — from healthcare to education, from productivity to mental health — is powered by a shared behavioural intelligence layer that truly understands the human it serves.
ModalityInputArchitectureStatus
LanguageTokensGPTClaudeLLaMA✓ Solved
VisionPixelsViTsCLIPDINOv2✓ Solved
Physical WorldActions / PhysicsJEPAGenie◐ Emerging
AudioSpectrogramsWhisperAudioPaLM✓ Solved
Human BehaviourBehavioural State VectorsLBM◉ We're building it
Ongoing Research
What we're working on.
Six core research streams, each addressing an unsolved problem in modelling human behaviour at scale.
01
Behavioural State Vectors
The 240-dimensional continuous representation of a human's current state — cognitive, emotional, biological, motivational — updating in real-time.
Active
02
Causal Behavioural Graphs
Directed causal graphs capturing why behaviour happens — 230M+ validated edges across 8 discovery methods.
Active
03
Episodic Memory Architecture
Multi-scale memory compressing 15-minute episodes and 15-year developmental arcs into a unified retrieval system.
In Progress
04
Self-Supervised Learning
7 self-supervised objectives for training on behavioural data without ground truth — next-token prediction for behaviour.
Active
05
Safe Reinforcement Learning
Mathematical safety guarantees for behavioural interventions. When should a model nudge, warn, or stay silent?
In Progress
06
Population Intelligence
Scaling individual models to population-level insights without Simpson's paradox — from n=1 to n=10M.
Forming
Architecture
How LBM works.
A five-layer architecture transforming raw signals into causal understanding and safe intervention.
LBM Architecture Stack
Each layer builds on the one below — from raw signal to actionable understanding.
Ingestion
Digital activitySleep patternsPhysiologyEnvironmentSocial signals
Episodes
Micro (15-30m)Macro (days-weeks)Context windowsState inference
BSV Engine
240-dim vectorReal-time updatesLatent spaceTemporal dynamics
Causal Graph
8 discovery methods230M+ edgesRoot-cause inferenceCounterfactuals
Intervention
Safe RL policyNudge / Warn / SilenceSafety boundsLong-horizon RL
Built by the Dotsin founding research team.
The Team
Who's building this.
A small, focused founding team with deep expertise across AI, behavioural science, genomics, and quantitative finance.
Publications & Partnerships
Our work.
Explore our latest research papers, code repositories, models, and industry collaborations.
Correlation Is Not Enough: Embedding Human Metadata for Individual Causal Discovery
Research Paper

Correlation Is Not Enough: Embedding Human Metadata for Individual Causal Discovery

This paper argues that correlation alone is too thin to support serious human modeling, because it collapses individual structure into population averages. It proposes embedding human metadata to recover person-level causal organization, enabling more grounded inference, personalization, and state estimation.

You Are in Control of Your State: Why Human Outcomes Are Controllable Through Causal State Intervention
Research Paper

You Are in Control of Your State: Why Human Outcomes Are Controllable Through Causal State Intervention

This paper argues that human outcomes are mediated by dynamic latent state at the moment of decision, not by observable inputs alone. It proposes a causal state-intervention framework that combines attentional bottlenecks, longitudinal behavior, and state weighting to define practical levers for change.

Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation
Research Paper

Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation

This paper proposes a Bayesian framework that uses genomic priors as a personalized anchor for physiological interpretation. It separates constitutional baseline from environment-driven deviation, allowing earlier and more individualized interpretation of physiological signals.

Human Modelling Requires a Causal Architecture of Behaviour and Biology, Not Correlation
Research Paper

Human Modelling Requires a Causal Architecture of Behaviour and Biology, Not Correlation

This paper argues that human modeling must be built on causal architecture rather than correlation-only representations. It frames behavior and biology as interacting regimes, which makes inference more structurally valid and more useful for intervention design.

Expert review: You Are in Control of Your State
Research Paper

Expert review: You Are in Control of Your State

This review validates the central claim that human outcomes are not fixed outputs but controllable trajectories shaped through causal intervention. It serves as external commentary supporting the state-based framing and strengthening the broader research program.

Code & Model

Dotsin GitHub

Our GitHub organization publishes the reproducible layers of the research stack, including code, model artifacts, benchmark tooling, and validation utilities. It functions as the public engineering surface of the lab, supporting transparency, reproducibility, and iterative development.

Benchmark repo ~ Biomedical BERT Embedding Quality & Inference Benchmark
Code & Model

Benchmark repo ~ Biomedical BERT Embedding Quality & Inference Benchmark

This benchmark demonstrates whether biomedical embeddings preserve meaningful semantic structure while remaining efficient enough for production inference. It evaluates similarity quality, hard-negative separation, and operational performance under realistic hardware constraints.

Code & Model

Hugging Face model

This model release exposes the retrieval layer of the biomedical inference stack as a controlled, production-oriented embedding model. It is designed to index text by semantic and causal proximity, making it a reusable foundation for downstream research workflows.

Intel Partnership
Partnership

Intel Partnership

We’re now an official Intel AI Partner, optimizing LBM inference on Xeon 6 with OpenVINO and IPEX for real-time behavioural intelligence at scale.

NVIDIA Partnership
Partnership

NVIDIA Partnership

Accepted into NVIDIA’s Inception Program, we’re building the behavioural AI infrastructure that makes human understanding as foundational as language intelligence.

The room is open.

We're building the first Behavioural Foundation Model from India — not as a company project, but as a national scientific effort.