Open to opportunities

Extracting a reliable signal from hidden states

Ten years asking one question across affect, patient risk, robot behaviour, genomes and multi-agent coordination — and making the resulting models trustworthy enough to deploy. Currently building plant genomic foundation models at Living Models.

One question, several settings

How do you extract a reliable signal about a hidden state — affect, patient risk, robot behaviour, genome function, agent coordination — from multimodal or sequential data, and make the resulting model trustworthy?

A hidden state is never observed directly. It is read through observable cues, each only partly valid — Brunswik's lens, 1952. The weakest cues often carry the information the salient ones do not, which is why subtle signals are worth the trouble of modelling.

Interpretability does more than explain. It lets you test what a model learned against the literature of the field, and sometimes find explanations deeper than the ones already described. Every step below is the same question applied to increasingly complex data.

Brunswik lens model: a hidden state is externalized as distal cues, perceived as proximal cues, and attributed as a perceptual judgment; automatic recognition and synthesis mirror this at the technical level.
The Brunswik lens (1952) — the framing behind every project on this page. Figure after Brunswik; technical-level annotation as used in affective computing.

Seven steps, one question

Apr 2026 — Present · step 7
AI Researcher — Plant Genomics
Living Models — Paris

Genomic foundation models (BOTANIC-0 / BOTANIC-1) across dozens of crop species: long-context Mamba/SSM and hybrid architectures at single-base resolution, used for questions such as identifying the role of a DNA sequence in hormone synthesis. Post-training goes beyond masked language modelling — functional supervision, regulatory atlases, breeding-relevant readouts — with evaluation suites meant to be meaningful to breeders rather than to leaderboards. The stack itself is agentic: Claude and Gemma orchestrate training runs and the deployment of an LLM augmented with bioinformatics tools.

The models are public: the Botanic1 report on Hugging Face, alongside the Botanic0 and Botanic1 weights (0.3B to 2B) and the pretraining dataset. The work is written up in the Botanic1 preprint.

A model factory rather than a model. The corpus spans 320 species across 102 families — the broadest taxonomic coverage of any plant genomic language model so far — and the backbone is a stack of bidirectional Mamba2 blocks released at four sizes from 318M to 3.2B parameters. Pre-training runs at 8,192 bp and extends to 131,072 bp, which is what makes needle-in-a-haystack retrieval across a genome possible at all. On the aggregate balanced score, all four Botanic1 sizes rank above every other model evaluated.

Four-panel figure. a, model training: 320 plant species and 59 Gbp of data feeding a BiMamba2 stack of 24 to 80 blocks predicting masked single-nucleotide tokens, with released sizes from 318M to 3.2B parameters. b, six evaluation regimes from frozen embeddings and LoRA to full fine-tuning, zero-shot scores, interpretability and long-context retrieval. c, context sizes from 512 bp probe tasks to 131,072 bp retrieval. d, model ranking by balanced score, with the four Botanic1 models occupying the top four positions ahead of PlantCAD2-L and PlantCaduceus.
The Botanic1 model factory: training, evaluation regimes, context sizes, and the test leaderboard. Figure 1 from Barozet, Cabeli, Ogier du Terrail, Rukhovich, Janssoone et al., bioRxiv 2026 (CC BY 4.0).
2025 — Present · ongoing
Independent Research — Multi-Agent Coordination
Dotomics · Concilium

Coordination between LLM agents inspired by hormonal mechanisms, designed explicitly with alignment in mind: long-horizon tasks, democratic voting, feedback signalling as a shared control channel. Documented an open-loop coordination failure — signals emitted but no longer feeding back into shared state — which is the kind of silent degradation a multi-agent system has to be built to notice.

Sep 2025 — Feb 2026 · step 6
AI Research Scientist — Medical Speech Analysis
e-sensia — Paris

Pathology detection from vocal biomarkers for emergency triage: a life-critical variant of the same problem. Fine-tuned audio models (Whisper, Wav2Vec2, EnCodec) alongside raw signal processing, fused into text, acoustic and multimodal risk scores under latency constraints — diarization, transcription, information validation, spoken response, with patient history in the loop. Led the ESYNAPSE grant, a France 2030 "Pionniers de l'IA" laureate worth up to €10M, and worked with clinical partners to validate models on real emergency medical data.

Pipeline diagram: patient voice enters diarization, which feeds transcription and acoustic pathology detection; transcription feeds information validation and text-based pathology detection; the text and acoustic risk scores meet at multimodal pathology detection, which also receives patient record and history from a database lookup; information validation drives a spoken response back to the patient.
The triage pipeline: two independent readings of the same call — what was said and how it sounded — fused into one risk score, with the patient record joining at the fusion step. Redrawn from the original architecture diagram.
Apr 2022 — Aug 2025 · step 5
Multimodal ML Expert — Social Robotics
Enchanted Tools — Paris

Perception → reasoning → action, in real time, under human-robot safety constraints for the Mirokai robot. Here interpretability stops being an abstract topic and becomes a deployment constraint.

Two Mirokai robots: a humanoid upper body on a rolling ball base, one pushing a cart and one carrying a tray.
Mirokai — an emotional service robot with an interactive head, opposable thumbs and a rolling-globe base.

Making a small model speak Mirokaï. The robot needs an LLM to track the history of an exchange, trigger the right tools and stay inside its narrative backstory (its "Lore") — but embedded hardware makes inference on a large model impractical. I generated synthetic dialogues with a larger model, cleaned the dataset of hallucinations and other errors, and used it for PEFT-LoRA fine-tuning of Llama-3.2-1B: a teacher-student route to large-model behaviour at an embeddable size.

Generating the data you cannot collect. A parameterised MetaHuman dataset with emotion, intensity, gender, ethnicity, environment, distance, viewing angle and robot size as controlled axes. Variations that are impossible to balance in real collection become dimensions you can sweep.

Synthetic MetaHuman dataset: graded surprise and anger intensity ramps, the same character rendered across four environments, and an adult-size versus child-size viewpoint comparison.
Parameterised MetaHuman dataset — emotion intensity ramps, environment variation, and the same expression seen from adult versus child robot height.

Then using it to anticipate deployment bias. Emotion recognition collapses when the robot is small — the same model that reaches 0.86 accuracy at adult size drops to 0.04 at baby size, a factor of twenty — and the predicted score drifts with age, gender and ethnicity.

The point is when this was measured: before deployment, not after. A synthetic dataset you control is the cheapest way to find out that your perception stack has a blind spot at the exact height a child would look up from.
Four panels. Two heatmaps of emotion recognition accuracy by robot size against distance and yaw angle, ranging from 0.038 at baby size to 0.95 at adult size. Two boxplots of predicted emotion score by age and gender, and by race and gender, both showing systematic drift.
Accuracy by robot size × distance and × viewing angle (top); predicted emotion score by age, gender and ethnicity (bottom). Measured on the synthetic dataset, pre-deployment.
2021 — 2022 · step 4
Postdoctoral Research Scientist — Human-Agent Rapport
Inria COML, Justine Cassell's team — Paris

Modelling rapport in conversational human-agent interaction — the closest thing to pure HCI in my career, and the direct continuation of the HHAI 2022 paper in the deep dive below. Interpreted dyadic models (LSTM/GRU with attention), moving from estimating rapport to generating behaviour that builds it, validated through perceptual experiments with humans. The question underneath: how does an agent's behaviour shape its user's?

Architecture diagram: databases of dyadic interactions and real-time interaction feed video and sound extraction, then a feature extractor using openFace, openSmile and ASR; a feature analyzer holds a conversational strategies classifier, a rapport estimator and an NLU module; a dyadic model with task manager and social reasoner drives generation modules — NLG, Beat, Smartbody/Greta and text-to-speech — rendering an animated agent; model evaluation and model explainer modules hang off the analyser.
The full loop the rapport estimator sits inside: extraction, analysis, a dyadic model, then generation and rendering — with the explainer wired in rather than bolted on afterwards.
2018 — 2021 · step 3
Research Scientist — Interpretability in Healthcare
Semeia — Paris

Predicting medication non-adherence from SNIIRAM reimbursement records on French National Health Data, taking prediction accuracy from 60% to 90%. SHAP and LIME were not a presentation layer here — they carried real clinical decisions, which is a different bar for an explanation. Published at NeurIPS ML4H and MICCAI.

Clinical interface in French: a banner states the patient has a medium initial risk of stopping treatment at three months and suggests closer follow-up; below, a force plot on a treatment-drop-out risk axis from 45 to 60 percent shows previous chemotherapy and age pushing the score up and long-term-condition seniority and metastatic status pulling it down, landing at 49 percent against a 53 percent average.
What the clinician actually saw: not a score but the reasons for it — each patient feature shown pushing the drop-out risk up or down from the population average. French, as deployed.
2014 — 2018 · steps 1–2
PhD — Multimodal Social Signals
Sorbonne Université / ISIR / Télécom ParisTech

Recognising social signals for affective virtual agents, with no black box anywhere in the pipeline. Five publications (ICMI, IVA, WACAI, RIA, LREC) and a visiting period at USC ICT.

What a social signal looks like to a machine. An interpersonal attitude is not observable, but the facial actions that carry it are. Brow raisers, lid tighteners, lip corner pullers — the Facial Action Coding System turns a face into a vocabulary of discrete, timestamped events that a rule miner can work with, alongside prosody and head motion.

Twelve labelled close-up crops of facial regions, each illustrating one Action Unit: inner and outer brow raiser, brow lowerer, upper lid raiser, cheek raiser, lid tightener, nose wrinkler, upper lip raiser, lip corner puller and depressor, chin raiser, lip stretcher.
The Action Units extracted from each speaker. Janssoone et al., LREC 2020 (CC BY-NC).

From corpus to interpretable rules. Those symbolised signals feed TITARL temporal association rules, then a discrimination metric isolates the rules specific to one attitude rather than the ones every speaker produces. What comes out is readable: a set of timed if-then statements, not a weight matrix.

Two-part diagram of the SMART pipeline: study corpus to signal extraction to simple rule computation to discrimination between common and specific rules; below, the temporal association rule mining loop of rule creation, division and refinement.
The SMART pipeline (top) and the TITARL rule-mining loop (bottom) — original thesis figures, labelled in French.

Two strategies, both put in front of users. Either clone a speaker's signals onto the agent directly, or build a model from the corpus and let the rules drive generation. Both routes end at the same place — human raters — and both feed back: if the clone is not faithful enough, the extraction gets adjusted; if the evaluation is inconclusive, the model does.

Flow diagram: data collection feeds social signal extraction, producing a set of signals; a cloning strategy reproduces the original data and a rule strategy builds a model from it, each yielding data evaluated by users, with dashed feedback arrows returning to extraction, to the model, and to data collection.
The cloning and rule-based strategies, with their evaluation loops. Janssoone et al., LREC 2020 (CC BY-NC).
The result: a speaker's social signals cloned onto a virtual agent, from the POTUS corpus of weekly presidential addresses (LREC 2020).

A detour, and the most cited thing I have written. Separate work at LIG Grenoble, on the other side of the same problem: instead of reading a human's gestures, guiding them. OctoPocus3D shows a whole 3D gesture set as coloured pipes starting from the user's hand, thinning as the recogniser rules candidates out. Concurrent feedback raised the recognition rate by 10% for novices, and showing only the upcoming portion of a gesture cut completion time by 8% — less visual clutter beat more anticipation.

Three frames of the OctoPocus3D guide on a black grid: first the full set of coloured 3D gesture pipes labelled with city names, then the set thinning as the user follows one gesture, then only the two most likely gestures remaining.
OctoPocus3D: the full gesture set (a), thinning as the user commits (b), then resolved (c). Delamare, Janssoone, Coutrix & Nigay, AVI 2016 · © ACM, author's version.

Three threads running through all of it

⚖️

AI Safety

Risk-prevention models where being wrong costs a life (Semeia, e-sensia). Alignment in social robotics and deployment bias anticipated before deployment rather than discovered after (Enchanted Tools). Multi-agent coordination designed around alignment from the start (Concilium).

🔍

Explainability

SHAP and LIME driving clinical decisions (Semeia), interpretable temporal rules instead of a black box (PhD), and interpretability as a hard deployment constraint on embedded robots (Enchanted Tools). An explanation is only worth something if it changes what you do next.

🧠

Human-AI Interaction

Rapport in human-agent interaction (Inria, HHAI 2022) and how an agent's behaviour shapes its user's — measured, then fed back into how the agent generates behaviour in the first place.

Estimating rapport in videoconference tutoring

Grimberg, Janssoone, Clavel & Cassell, HHAI 2022 — ENS Paris · Carnegie Mellon & Inria · Télécom Paris. Teenagers tutor each other in algebra over Skype, two sessions of about thirty minutes per dyad. The corpus comes from Madaio et al. (2017); we did not collect it, we re-exploited and re-annotated it.

14
usable dyads
2,209
30 s slices with a valid rapport annotation
1,588
slices where every feature is extractable
α = 0.61
Krippendorff inter-rater agreement

Annotating a subjective perception

The target does not exist in the signal; it has to be constructed. Four independent judges recruited on Amazon Mechanical Turk rated every 30-second slice on a 1–7 Likert scale. The most distant rating was discarded and a bias correction applied to the remaining three.

α = 0.61 is moderate agreement. In other words, even among humans, rapport is not a consensus — which is an implicit performance ceiling for any model trained on these labels, and a good reason to distrust MAE comparisons at the third decimal.

Features and protocol

High-level features from OpenFace (non-verbal) and openSMILE (para-verbal), in two families. Individual features describe one participant — SpeechRate, F0, Loudness, VoiceProb, PosFace, GazeChange, HeadRotation, each as mean, standard deviation, skew and max. Relative features describe the two participants against each other — RelSpeechRate, RelF0, RelLoudness, RelPosFace. Notation: t = tutee, T = tutor.

The split is by dyad, not by slice, so no participant leaks between train and test: 10 dyads / 1,546 slices for training, 2 dyads / 358 for validation, 2 dyads / 305 for test. Hyperparameters were tuned on validation, then evaluated once on test — no going back and forth.

Results — 5 models × 3 feature sets

MAE on the test set, lower is better. Rapport is annotated on a 1–7 scale.
Model Individual Relative All
Support Vector Regressor1.0921.1871.232
Random Forest Regressor1.0781.2201.074
Decision Tree Regressor1.1551.2561.155
Gradient Boosting Regressor1.0721.1261.086
MLP Regressor1.2271.2171.275
Dummy regressor (predicts mean rapport)1.293

Every model beats the baseline, but not by much. Gradient Boosting on individual features (1.072) and Random Forest on all features (1.074) are separated by a gap that means nothing.

What SHAP says

SHAP beeswarm plot for the model trained on individual features. tLoudnessStd and TLoudnessStd rank first and second, followed by TPosFaceMean and tPosFaceMean.
Individual features. Vocal intensity variability dominates for tutee and tutor alike: a monotone delivery drags the estimate down. Positive facial expression follows, with the tutor's effect more pronounced — consistent with the tutor driving the interaction.
SHAP beeswarm plot for the model trained on all features. RelF0Skew ranks third and RelPosFaceStd fifth, ahead of several individual features.
All features. The interesting result: RelF0Skew comes 3rd, ahead of TPosFaceMean, and RelPosFaceStd (5th) outranks tPosFaceMean. How the two voices and the two faces move relative to each other carries signal the isolated expressions do not.

Individual features still hold most of the top ranks — synchrony counts without dominating. But it counts enough that ignoring the dyad loses information.

Why this is useful

This is where interpretability pays off, because a score alone does not tell you what to do. "Rapport in this slice is 3.7" is measurable and not actionable. "Synchrony of positive expressions raises rapport" is an instruction for an agent's behaviour-generation module: if synchrony on a feature raises rapport, it can be built in as an objective. And nothing in the pipeline is specific to rapport — the same chain applies to intimacy, trust, or any other annotatable conversational phenomenon.

Owned limitations

Weak predictive power

1.072 against a 1.293 baseline on a 1–7 scale, with statistical power under 50%. Better than chance, not much better, and we never presented it otherwise.

A trade-off on missing data

Audio extraction fails on part of the slices. Excluding them would hurt performance on an already small set, so they were kept — at the cost of an imbalance between acoustic and visual features.

Moderate inter-rater agreement

α = 0.61 means the target itself is noisy. Comparing MAEs at the third decimal on these labels would be meaningless.

14 dyads, a single context

Algebra tutoring between teenagers over videoconference. Nothing guarantees the feature ranking holds anywhere else.

What I would do differently today

A larger, more diverse corpus

Several contexts and cultures, to test whether the feature ranking generalises at all.

Sequential models

Rapport builds over time; treating each 30-second slice in isolation denies that. This is exactly what the Inria postdoc went on to do.

A causal design

Manipulate an agent's expression and measure the effect, rather than observing correlations.

Natural-language explanations

Translate SHAP into instructions a designer can read, not scatter plots they have to interpret.

Research outputs

10+
Years in ML Research
~121
Citations (Scholar)
6
h-index
7
Selected Publications
View all on Google Scholar → Résumé (PDF) PhD summary (PDF)

Let's build something meaningful

Interested in AI safety, interpretability, or human-AI interaction? I'm always open to conversations about research collaborations or opportunities.