Reading an imagined object out of EEG, and the number that nearly buried it

A system that turns an imagined object into 3D, in real time. It ran end to end, returned chance, and then the actual result showed up disguised as a failure.

The number was 0.20. Chance was 0.25.

Everything around it worked. The cap streamed 64 channels over LSL, the encoder returned a vector, the catalogue search found a photograph, the generator turned that photograph into a mesh, and the mesh arrived inside a Meta Quest 3 as a cloud of particles that slowly gathered into a shape while the participant sat there with conductive gel in their hair. End to end. In real time. With a real person in the chair.

The infrastructure was fine. The decoding was noise.

That was the low point, and it is the part worth writing about, because the result that ended up carrying the thesis was already sitting in the same data. It was answering a question we had asked wrong.

The problem, briefly

Reading out what a person is looking at from their brain activity is close to solved if you have an fMRI scanner. MindEye identifies the viewed image with 93.8% accuracy, MindEye2 with 92.1%. The catch is the scanner: a room, a magnet, and a subject who has to stay still.

EEG is a cap. It costs a fraction as much, it goes anywhere, and it has millisecond timing. It also has poor spatial resolution and a noise floor that makes the fMRI numbers read like fiction. On THINGS-EEG2, NICE-EEG gets 15.6% and ATM 21.1% against a chance level of 0.5%. Given the physics that is respectable, and a long way from useful.

Two things were true of all of that work. It stops at flat 2D images, and it aims at the same target, CLIP.

Silvia Bracco and I wanted something slightly different: EEG recorded while a person imagines an object, with nothing in front of them, turned into a three-dimensional thing you can walk around. That was our master’s thesis at Politecnico di Torino, supervised by Andrea Sanna and Federico Manuri. We called it DreamWalking.

The one decision the rest of it hangs on

Hunyuan3D 2.1 produces a shape in two or three seconds, and it conditions on DINOv2 features.

That sentence is the architecture. Align EEG to CLIP like everybody else and you need a translator in the middle to reach a generator. Align EEG to DINOv2 instead and the generator already speaks the language you trained the encoder to speak. One alignment objective buys semantic proximity and generation quality at the same time.

It looks obvious written down. It did not look obvious at the time.

Never ask a noisy vector to invent a shape

The second decision came out of accepting how bad the signal is.

The naive pipeline hands the predicted embedding straight to the generator and asks it to build something. That works only if the embedding is close to correct, which an EEG-derived embedding is not.

So we put a search in the middle. Roughly 14 million photographs from ImageNet-21k, indexed by their DINOv2 embedding, searched with FAISS in under 30 milliseconds on a CPU. The estimated vector no longer has to specify a shape. It only has to land in the right neighbourhood, and the search snaps it onto an object that actually exists.

The side effect turned out to matter as much as the robustness. Retrieval makes the system legible. You can look at the photograph it picked and see what it thought you were imagining, before a single triangle gets generated. When it is wrong you know immediately, and you know how it was wrong.

Block diagram of the DreamWalking pipeline, from EEG acquisition through the channel-agnostic encoder, alignment to DINOv2, FAISS search over ImageNet-21k, 3D generation and volumetric rendering, with a red return path feeding user confirmation back into the LoRA trainer and the catalogue.
ARCHITECTURE Six blocks forward, from the amplifier to the object in the headset. The red path is the return: what the user says goes back into their personal adapter and into the catalogue. Tap to enlarge.

Uncertainty you can look at

The encoder does not produce an object. It produces a distribution over plausible objects, and the spread of that distribution is the system telling you how sure it is. Materialise it as a clean mesh and you throw the spread away and replace it with a confident wrong answer.

So the object arrives as a Niagara particle cloud in Unreal Engine 5, advected by curl noise. Density and coherence encode confidence. An unsure guess stays a drifting fog. A confirmed one consolidates into a solid surface.

A rubber duck rendered as a loose cloud of glowing yellow particles against a starfield, its outline recognisable but its surface unresolved.The same rubber duck as a solid shaded mesh on the same starfield, with a defined beak and eye.
DREAM AND COMMIT One object, two states. First the hypothesis, while the system is still unsure of it. Then the mesh, after the participant said the word out loud.

I pitched this as an aesthetic choice. It was the only honest way to display the output, and I did not work that out until much later.

Confirmation happens by speaking. The participant says the word for what they were imagining, the speech goes to text, the text goes to a text-to-image model that produces the truth image, and the cosine similarity between truth and hypothesis decides whether the cloud solidifies or dissolves and retries. Either way the pair lands in a replay buffer, and a rank-16 LoRA adapter updates on it in one to three seconds. Saying the word both answers the system and teaches it.

An hour in the chair

The setup was not gentle. Sixty-four active electrodes on the extended 10-20 montage, gelled one at a time, then a Quest 3 strapped on top of the cap, then over-ear headphones. One continuous session of about an hour.

We had no trigger hardware, so we wrote our own acquisition client against the g.NEEDaccess API and timestamped every event coming out of the VR scene against the acquisition clock, rather than against bytes written to disk, which lag behind by a roughly constant amount. Inference ran on an H200 node of the Politecnico HPC cluster, reached over a VPN tunnel from the workstation sitting next to the participant.

The DreamWalking operator console during a session. A left column lists pipeline states including VPN connecting, SLURM queued, SLURM running and tunnel open. The centre shows the retrieved hypothesis image of a rubber duck. The right panel renders the generated 3D mesh, labelled retrieved: papera, similarity 0.870.
STUDIO The operator console mid-session. Left, the state machine, tunnel and SLURM queue included. Centre, the photograph the search returned. Right, the mesh built from it. Tap to enlarge.

The virtual scene was deliberately boring: a dark starfield, a reflective floor, four broken stone columns, one small light drifting slowly across them. Anything more interesting would have painted evoked visual activity all over the imagery signal we were trying to read. Press A, the world goes black, imagine the object. The world comes back, the columns reassemble, the cloud appears. Say the word.

Andrea, our supervisor, cut his hair so we could test the cap on him. It is in the acknowledgements, and it is the detail I tell people first.

The number that looked like a refutation

Back to 0.20 against a chance of 0.25.

The production pipeline was at chance. Worse, when we ran the scientific test, the one meant to establish the hypothesis the whole architecture rests on, we got 0.322 where chance is 0.500.

Below chance. On the central claim.

There is a specific feeling attached to a number like that. It is not disappointment. It is the suspicion that you have built an elaborate machine on top of something that was never true.

The hypothesis was this: the brain representation of an imagined object is linearly mappable into the DINOv2 CLS space, and a linear map fitted on some concepts will place a concept it has never seen in the right region of that space. If that holds, retrieving unseen objects from a catalogue makes sense. If it does not, nothing downstream is justified.

We had tested it the obvious way. Hold out one concept, fit ridge regression on the rest, check whether the held-out concept lands where it should against a gallery containing all the training concepts.

The reversal

Ridge regression contracts its predictions toward the centroid of the training set. That is what the penalty term is for. It is not a bug.

Now hold out one concept. By construction that concept is the most peripheral point in the space, the one furthest from the centroid everything is being pulled toward, so it is systematically the most anti-correlated with its own prediction. The protocol was not measuring the hypothesis. It was measuring the regulariser, and it returned the exact inverse of the truth.

The fix is to hold out both concepts in every comparison. Leave-pair-out instead of leave-one-out. Neither term is then the centroid of the training distribution, the bias applies equally to both sides, and it cancels by symmetry.

Same data. Same map. Same features.

0.744 as the per-concept mean, 0.724 trial by trial, against a chance of 0.500 and a permutation null distribution sitting at 0.494. p = 0.005 over 200 permutations.

Bar chart titled zero-shot evaluation protocols compared. Leave-one-out with a full gallery reaches 0.322, below the dashed chance line at 0.50. Leave-pair-out, the canonical test, reaches 0.744.
THE REVERSAL Hold out one concept and it is always the point the regulariser pulls hardest away from, which is what drags the left bar under chance. Hold out both and the pull cancels.

None of the concepts in that test were seen during training. A linear map fitted on some concepts puts a new one in the right region of DINOv2 space, well enough to tell it from a distractor. That property is what the retrieval architecture needs in order to make sense, and it is the one thing here I am confident about.

The features behind that number are unglamorous, which is the best thing about them. Multiband spatial covariance at the scalp, 31 channels, and a ridge regression. No inverse model, no anatomical template, nothing that has to be fitted per subject before the test can run. Anybody with a scalp recording and an afternoon can try to break it.

What it does not say

Two-way identification is a choice between two alternatives. The catalogue has 14 million entries. Top-k retrieval over the full catalogue was never evaluated, and the distance between “tells A from B” and “finds the right concept among 21,000 categories” is enormous. That is the first missing test, and I would rather say it than have it extracted from me.

The decoding is also categorical, not fine. Four concepts spread across different categories (dog, circle, cup, scissors) decode at up to 0.787. Four objects from within one category (chair, cup, watch, scissors) sit at 0.267 and 0.296 depending on the feature space, which is chance. Our own sessions showed the same shape: apple versus aeroplane came out at 0.756 with p = 0.003, replicated across two independent signal representations, while four-way decoding over the full set stayed at chance.

EEG imagery tells you what kind of thing someone is imagining. Not which one.

And the main test is one subject.

What happens next

The encoder is the piece that gets replaced. Not patched: replaced, by the direct linear map from covariance features to DINOv2 CLS, refitted on imagery data instead of on passive viewing. The rest of the pipeline already runs.

We know how much the current one costs us, because we measured it rather than guessed. Push the same imagery data through the production path, frozen encoder and LoRA included, and it still identifies the right concept out of ten at 0.248 against a chance of 0.100. A decoder built for that task, on the same data, gets 0.276. The encoder does not destroy the signal. It hands over a version of it degraded by a margin wide enough to sink retrieval and narrow enough to be worth recovering, which is exactly the kind of failure you want, because it tells you what to build next.

The consequence I actually care about is smaller and more practical. If the decoding is categorical, a handful of well-chosen anchor concepts might be enough to fit the map for a new person. Five or ten minutes in the chair instead of a training corpus. That is the next thing to test, and it is testable now.

But the piece I keep turning over is the 0.322. It sat in a notebook and on a slide, and for a while it was a good reason to believe the whole idea was wrong. The data never changed. The question we put to it did.