Spatial Intelligence

Yimeng Liu

Yimeng Liu — Ph.D. Candidate in Computer Science and Engineering at Michigan State University, working on Spatial Intelligence

Ph.D. Candidate Computer Science and Engineering Michigan State University

I work on AI that has to know where it is — building a working understanding of a physical space out of incomplete sensors, and keeping that understanding current enough to act on.

The problem
  1. Sense
  2. Represent
  3. Model
  4. Reason
  5. Act
See the three scales

07 The record

The record, in full

Publications, education, service and current activity. Deliberately the last thing on the page: it should be one click away, never the first impression.

Publications

10 papers · h-index 6 · 124 citations counted on Google Scholar, 2026-10-02

10 papers

  1. Bioinspir.2022 cited 27

    Detection of Passageways in Natural Foliage Using Biomimetic Sonar

    Rongkun Wang, Yimeng Liu, Rolf Müller

    Bioinspiration & Biomimetics, 2022

    What I take from itMy first attempt at the question I am still working on: can structure in a space be recovered from returns that are sparse and indirect? Echolocation through dense foliage, where optical sensing simply fails. I would not have called this Spatial Intelligence then, but it set the habit of asking what a measurement can and cannot support.

  2. MobiSys'232023 cited 19

    mmLeaf: Versatile Leaf Wetness Detection via mmWave Sensing

    Maolin Gan, Yimeng Liu, Li Liu, Cheng Wu, Younsuk Dong, Huacheng Zeng, Zhichao Cao

    Proceedings of the 21st Annual International Conference on Mobile Systems and Applications, 2023

    What I take from itMIMO SAR imaging of a leaf at varying working range. This is where the RF modality entered the work — and where I first ran into the fact that a camera and a radar do not observe the same thing. Everything published after this is a consequence of that.

  3. MobiCom'242024 cited 26

    Hydra: Accurate Multi-Modal Leaf Wetness Sensing with mm-Wave and Camera Fusion

    Yimeng Liu, Maolin Gan, Huacheng Zeng, Li Liu, Younsuk Dong, Zhichao Cao

    Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, 2024

    What I take from itThe paper that made the problem explicit. Two modalities chosen because they fail under *different* conditions, fused so that perception survives losing either one. It is also where I stopped treating fusion as the interesting part: a fusion model is really a representation problem, and the representation is what has to survive modality loss.

  4. INFOCOM'252025 cited 18

    AeroEcho: Towards Agricultural Low-Power Wide-Area Backscatter with Aerial Excitation Source

    Yidong Ren, Gen Li, Yimeng Liu, Younsuk Dong, Zhichao Cao

    IEEE INFOCOM 2025 — IEEE Conference on Computer Communications, 2025

    What I take from itPassive agricultural sensing over wide areas, excited from the air. This is the piece that moved the question from single devices to sensing *infrastructure* — and it is the first time my own work depended on observers I do not control, which is a different kind of problem from owning a sensor.

  5. IMWUT'262026 cited 5

    Soilnutri: A Passive Metasurface-Based, Low-Cost System for Soil Moisture and Nitrogen Monitoring

    Juexing Wang, Binbin Xie, Yimeng Liu, Minhao Cui, Xiao Zhang, Guangjing Wang, Ke Sun, Zhichao Cao, Huacheng Zeng, Qingxu Jin, Younsuk Dong, Hui Li, Jie Xiong, Tianxing Li

    Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2026

    What I take from itPassive, low-cost, metasurface-based sensing of soil moisture and nitrogen. Environmental monitoring at the scale where the sensing substrate itself has to be close to free — the regime where "many cheap unreliable observers" stops being a thought experiment and starts being the only affordable design.

  6. SenSys'252025 cited 15

    Proteus: Enhanced mmWave Leaf Wetness Detection with Cross-Modality Knowledge Transfer

    Yimeng Liu, Maolin Gan, Huacheng Zeng, Yidong Ren, Gen Li, Jingkai Lin, Younsuk Dong, Xiaobo Tan, Zhichao Cao

    Proceedings of the 23rd ACM Conference on Embedded Networked Sensor Systems, 2025

    What I take from itCross-modality knowledge transfer: an RGB teacher guiding an mmWave student. The teacher was never the point. What mattered was recovering physical structure from sparse apertures — and that reframed the task as an inverse problem rather than an image enhancement problem, which is a claim about evidence, not about pixels.

  7. INFOCOM'252025 cited 13

    Adonis: Neural-Enhanced Fine-Grained Leaf Wetness Sensing with Efficient mmWave Imaging

    Yimeng Liu*, Maolin Gan*, Gen Li, Younsuk Dong, Zhichao Cao

    IEEE INFOCOM 2025 — IEEE Conference on Computer Communications, 2025

    What I take from itThe same physical state, represented compactly enough to be deployable rather than only demonstrable — range-Doppler, range-phase and range-azimuth maps with learned features and calibration. Proteus asked whether the state could be recovered; Adonis asked whether the representation survives contact with a real budget.

  8. CEA'262026 cited 1

    Comparative Analysis of Machine Learning Models to Restore Gaps in Multivariate Time Series Leaf Wetness Sensor Data

    Shivani Rana, Nawab Ali, Zhichao Cao, Jill C. Check, Martin I. Chilvers, Jaime Willbur, Benjamin Werling, Yimeng Liu, Younsuk Dong

    Computers and Electronics in Agriculture, 2026

    What I take from itThe unglamorous half of the thesis. A state you hold across time is only as good as what happens when your sensors stop reporting — and this paper is an honest measurement of how badly standard imputation fails on exactly the signal I care about.

  9. arXiv'262026

    AIoT-Based Continuous, Contextualized, and Explainable Driving Assessment for Older Adults

    Yimeng Liu, Fangwei Zhang, Maolin Gan, Jialuo Du, Jingkai Lin, Yawen Wang, Fei Sun, Honglei Chen, Linda Hill, Ruofeng Liu, Tianxing Li, Zhichao Cao

    arXiv preprint arXiv:2603.00691, 2026

    What I take from itEveryday driving as a continuous functional record instead of an occasional clinic snapshot — separating age-related change from situational factors like traffic and weather. This one is a design-principles and research-opportunities document: it proposes the sensing, modelling and context architecture, and states plainly that a reliable real-world monitoring system still has to be built and evaluated. The action stage of my agenda is genuinely unfinished, and this paper is an argument for how to finish it, not a result.

  10. arXiv'252025

    Hydra-Bench: A Benchmark for Multi-Modal Leaf Wetness Sensing

    Yimeng Liu, Maolin Gan, Yidong Ren, Gen Li, Jingkai Lin, Younsuk Dong, Zhichao Cao

    arXiv preprint arXiv:2507.22685, 2025

    What I take from itSix months of synchronised mmWave raw, SAR and RGB data, released publicly. A multi-modal claim is not checkable from prose alone, so this was an attempt to make ours checkable by someone who is not us — and it is the first step toward Question II, what can be known together.

Complete profile on Google Scholar

Education

  • 2023 — Ph.D. candidate, Computer Science and EngineeringMichigan State University — Advised by Dr. Zhichao Cao
  • 2019 — 2023 B.S., MathematicsVirginia Tech — Undergraduate research with Dr. Rolf Müller

Positions

  • Summer 2022 Research intern, EIN LabMichigan State University
  • 2023 — Ph.D. researcher, Spatial IntelligenceCSE, Michigan State University

Service

Invited reviewer

  • UbiComp / ISWC · UbiSense 2025
  • IEEE Internet of Things Journal
  • IEEE Transactions on Mobile Computing
  • IEEE Internet Computing
  • HAI 2025 · Poster track

Program committee

  • IEEE DSAA 2025

Talks & conferences

  • 2024 ParticipantNSF Workshop on Sustainable Computing, Purdue University
  • 2024 Presented HydraACM MobiCom, Washington, D.C.
  • 2025 Presented ProteusACM SenSys, Irvine, CA

Invited talk

Invited talk, University of Hawai'i at Mānoa — October 2024

Updates

  • Jul 2025 The Hydra-Bench multimodal leaf-wetness dataset (mmWave raw + SAR + RGB) is released on arXiv.
  • Jul 2025 My journey from math to computing was highlighted by MSUTODAY.
  • Mar 2025 Proteus was accepted at ACM SenSys 2025.
  • Dec 2024 Adonis and AeroEcho were accepted at IEEE INFOCOM 2025.
  • Oct 2024 Invited talk at the University of Hawaiʻi at Mānoa.
  • Aug 2024 Hydra was accepted at ACM MobiCom 2024 and presented in Washington, D.C.

Full timeline

Portrait of Yimeng Liu
Yimeng Liu
Ph.D. Candidate, Michigan State University
428 S Shaw Ln, East Lansing, MI 48824

Everything above is verifiable in the CV. If you are reviewing a submission or a collaboration request, that is the shortest path.

01 Spatial Intelligence

What I mean
by that term

Every problem I have published is the same question at a different scale: how much of a space's physical state can be recovered from returns that are each partial, each biased, and each blind in its own direction?

Spatial Intelligence

The ability of a system to hold a working model of a place — what is there, how it is arranged, how it is changing, how uncertain the model is, and what should be done about it — built from observations that no single sensor could ever supply on its own.

The same question at three scales Three panels, each a published problem at a different scale, each pairing what an optical channel can recover with what a penetrating channel recovers instead. A leaf: optical reads the surface, millimetre-wave returns scatter off internal water layers. Soil: no camera survives, a passive metasurface reports moisture and nitrogen. A space: an optical camera loses a pedestrian behind an opaque obstruction while a radio return still resolves position. All three produce a partial, biased estimate of physical state. MILLIMETRE A leaf Hydra · Proteus · Adonis · mmLeaf OPTICALsurface only mmWAVEinternal structure CENTIMETRE Soil PASSIVE SURFACE Soilnutri OPTICALcannot be deployed PASSIVE RFmoisture, nitrogen METRE → KILOMETRE A space OCCLUDER Driving · Biomimetic sonar · AeroEcho OPTICALlost behind occlusion RADIOposition, not identity WHAT ALL THREE PRODUCE a partial, biased estimate of physical state — the object the rest of this work has to make trustworthy
  1. 01

    No sensor returns the world.

    A camera reads a surface. A radio return reads distance but not colour. A passive soil sensor reads moisture where nothing else can be deployed. Each is authoritative inside its own blind spot and useless outside it.

  2. 02

    The blind spots are structured, not random.

    Occlusion, penetration depth and line of sight are geometry, not luck. That makes the question of what is knowable from a given set of returns an observability question — one you can reason about before you collect any data.

  3. 03

    A state you cannot trust is still a state.

    The useful object is not a picture of the space. It is a persistent belief about the space that carries its own uncertainty, survives the failure of any single input, and can be acted on.

The machinery is the same at every scale

Change the scale and you change the sensor, the physics and the deployment cost. You do not change the shape of the problem: partial and differently-biased returns, a physical hypothesis that has to explain them, a belief that has to survive being wrong, and a decision that has to be taken from something you are not certain about. That is what I mean by Spatial Intelligence — not a subject area, but one problem recognised at three magnifications.

How I work on it The rest of this page takes that problem apart: what the work concluded, how the capability is layered, what has actually been built, where it is going, and what I am trying to build at the end of it.

01 What I think

What the work
convinced me of

These are the three conclusions my own systems kept forcing on me — not a manifesto, and not three things I decided in advance.

  1. A sensor never returns the world.

    Every return is a partial, biased slice. A camera fails in the dark and in rain; a radio echo knows distance but not colour. Hydra only worked once I stopped treating mmWave and camera as two views of one truth and started treating them as two questions that fail differently — which is why this is a representation problem, not a fusion problem.

  2. Understanding has to persist.

    A snapshot is not understanding. Proteus and Adonis could both read a leaf well at a given moment; what neither could do is hold the same entity across time, notice what moved, what was occluded, and what is still unknown. Holding state long enough to reason over it turned out to be the harder half of the problem.

  3. The test is a decision.

    A representation earns its keep only when it changes what a system does. I care about models whose output is a grounded action — yield, slow down, alert, hand control back. Prediction that leaves the decision unchanged is a demo, and I would rather ship the slower version that actually moves something.

03 How I get there

Five things a system
has to be able to do

Five layers. The first two are carried by published work; the last three are where the agenda currently lives, and the site says so rather than implying otherwise.

  1. 01 Sense Published

    Read physical state through partial views

    Different sensors fail in different ways. The capability I have built and published is to make perception survive that fact: two modalities that cover each other's blind spots, and RF that recovers physical detail when light does not.

    Capability

    • Modality redundancy — perception that holds when one channel degrades
    • RF imaging that resolves state optics cannot, at longer working ranges
    • Sensing infrastructure for environments where cameras, power, or coverage fail

    Evidence

    • Hydra ACM MobiCom '24 mmWave depth + RGB fusion, transformer encoder over fused feature maps
    • mmLeaf ACM MobiSys '23 FMCW radar with MIMO SAR imaging of leaves at varying distance
    • AeroEcho IEEE INFOCOM '25 Aerial excitation and backscatter tags for wide-area agricultural sensing
    • Biomimetic sonar Bioinspir. Biomim. '22 Passive echo sensing for passageways in dense foliage
  2. 02 Represent Published

    Turn returns into stable spatial structure

    Raw signal is not a representation. This layer is about turning partial, noisy, modality-specific returns into structure that can be reasoned over: cross-modal transfer when a modality is missing, and calibrated imaging that stays precise across distances and conditions.

    Capability

    • Cross-modality transfer — keep sensing when the second modality is unavailable
    • Multi-view signal processing — separate measurements that carry complementary detail
    • Calibration that holds accuracy across environments instead of one lab setting

    Evidence

    • Proteus ACM SenSys '25 RGB teacher guides an mmWave SAR student; speckle reduction and phase-angle features
    • Adonis IEEE INFOCOM '25 Range-Doppler, range-phase and range-azimuth maps with learned features and calibration
    • Hydra-Bench arXiv '25 Six months of synchronised mmWave raw, SAR and RGB data as a public benchmark
  3. 03 Maintain In progress

    Hold the same entities across time

    The layer I am working towards. Representation becomes understanding only when identity, position and state persist across frames and across occlusions, with explicit record of what is still missing. I have no published result here yet; it is the next thing to build.

    Capability

    • Entity persistence — the same object recognised across time and modality
    • Occlusion and missingness as first-class state, not as gaps to be smoothed away
    • State that a downstream reasoner can query, not just a feature vector

    Evidence

    Nothing published here yet. This layer is the work, not the citation.

  4. 04 Reason Open question

    Make relations, uncertainty and futures explicit

    Also open. The question is not whether a network can regress a trajectory, but whether the system can state who is interacting with whom, how confident it is, and which futures it is choosing between.

    Capability

    • Entity relations — crossing, following, occluding — as objects, not as features
    • Calibrated uncertainty attached to each claim, so risk can be priced
    • Prediction evaluated as decision quality, not only as trajectory error

    Evidence

    Nothing published here yet. This layer is the work, not the citation.

  5. 05 Act Ongoing

    Let spatial understanding change what a system does

    My current, human-centred direction: continuous in-car sensing and simulator assessment for older drivers, where an early, gentle intervention matters more than a late certain one. It is ongoing research, and it is where I intend to close the loop from understanding to action.

    Capability

    • Continuous monitoring that flags risk earlier than a clinical threshold would
    • Simulator protocols that isolate the cognitive demands real roads impose
    • Community education, so insight actually reaches the people it is about

    Evidence

    • Cognitive Driving Ongoing project Senior driving safety and cognitive health — MSU, with Social Work and Epidemiology

04 What I have built

What this looks like
once it is built

Not a paper list. Each entry answers the same five questions, so the argument stays visible underneath the results.

Hydra testbed with millimetre-wave radar and camera observing a plant
Hydra testbed — radar and camera observing the same surface.

Sense — modality redundancy

ACM MobiCom2024

Hydra

Complementary mmWave and vision, so perception survives either one failing.

Problem
Leaf wetness duration drives plant disease, but the instruments that measure it are the least reliable number on a farm: they measure synthetic leaves, mis-time real ones by up to half an hour, and break on light, wind, or an unfamiliar plant.
Insight
A camera and a millimetre-wave radar fail in opposite conditions. Rather than deciding which one to trust, fuse them at the feature level so the pair keeps working when either degrades.
Built
A CNN that selectively fuses several mmWave depth images with an RGB frame into multiple feature images; a transformer encoder that relates those feature images into a single feature map; a classifier; training-time augmentation for generalisation. FMCW radar, 76–81 GHz.
Evidence
Up to 96% wetness-classification accuracy across varying scenarios. Deployed on the farm — including rainy, dawn and poorly lit nights — accuracy remains around 90%.
Why it matters
Published proof that multimodal sensing is a redundancy problem: the system stays useful when one of its eyes goes dark.
Proteus millimetre-wave radar testbed observing a plant
Proteus testbed — one modality learning from the other.

Represent — cross-modal transfer

ACM SenSys2025

Proteus

Keep sensing when the second modality is simply gone.

Problem
mmWave suits leaf wetness because it responds to subtle surface change, yet existing systems fuse it with an RGB camera — so they stop working in the dark, and their accuracy drifts with environment because SAR imaging is not well understood.
Insight
What has already been learned in the RGB domain can be transferred into the mmWave SAR domain. And the radar image can be made more informative before any learning happens at all.
Built
A noise-reduction algorithm that suppresses speckle, phase-angle data added to enrich surface texture at higher resolution, and an RGB teacher guiding an mmWave SAR student through cross-modality knowledge transfer. Prototyped on commercial off-the-shelf radar.
Evidence
Up to 96.3% accuracy across varied environmental scenarios, outperforming state-of-the-art methods.
Why it matters
Representation has to survive the loss of a modality. Proteus is my published statement of that requirement, and the point where a project became a thesis.
Adonis millimetre-wave imaging setup with radar, antennas and a plant target
Adonis setup — radar, scan geometry and a plant target.

Represent — calibrated precision

IEEE INFOCOM2025

Adonis

Three signal views and explicit calibration, so precision survives the environment.

Problem
Fine-grained physical state is what a downstream model is supposed to reason over, but measuring it precisely means resolving detail a radar returns noisily — and any system that is precise in a lab tends to stop being precise outdoors.
Insight
Range-Doppler, range-phase angle and range-azimuth views carry genuinely complementary information about the same surface. Using all three, with learned features and an explicit calibration stage, is what makes the measurement hold outside the lab.
Built
Three signal-processing mappings fused into one input; contrastive learning with a CNN that isolates wetness variation from environmental change; a wetness-extraction layer; and model calibration using transmitted and reflected signal strength. FMCW radar, 77–81 GHz.
Evidence
Leaf wetness MAE of 4.43 in controlled environments and 6.49 on a real farm, against 11.84 indoors and 14.32 outdoors for traditional leaf wetness sensors. Removing any one of the three signal views pushes MAE to 8–10; keeping only one pushes it to 14–17.
Why it matters
Measurement quality is a modelling problem, not an engineering detail. A world model is only ever as good as the structure handed to it.
Road scene used to illustrate the cognitive driving project
The driving space as an instrumented, continuously observed environment.

Act — grounded, human-centred

Ongoing projectIn progress

Cognitive Driving

Where the chain closes: a spatial state that changes what a car does.

Problem
Mild cognitive change raises driving risk well before it is visible. By the time behaviour looks impaired, risky driving has already begun — and most tools either flag clearly impaired drivers or hand the question to a costly clinical assessment.
Insight
There is an in-between stage: driving still largely intact, cognition already changing. It is observable continuously, if you instrument the vehicle and measure it against simulator-based capability assessment.
Built
Simulator protocols that isolate the cognitive demands of driving — cognitive load and attention; an on-board sensing system that tracks driving behaviour over time and surfaces early risk; realistic testing against road conditions; and community education so insight reaches the people it is about.
Evidence
Ongoing research. No published spatial-reasoning model or closed-loop policy is claimed here; the recruitment flyer describes the study as it currently stands.
Why it matters
This is the layer I intend to close the loop on. Acting on a physical world is only defensible when the state is good enough to act on — and when the people affected are part of the design.

05 Where this goes

Three questions I am
working towards

The published work above answers parts of two of these. The third is the one I am spending the most time on, and none of it is finished.

  1. I In progress

    What can be known from physical evidence?

    Observability, not just accuracy

    Sparse measurements rarely admit a single explanation. A handful of radio returns can be produced by many different geometries, and an image that looks right is not the same as a surface that would actually scatter that signal. I want inference that commits to an explicit physical hypothesis and then tests whether the hypothesis explains the measurement — instead of a model that produces a plausible picture nobody can check.

    Published work covers the sensing half. The inverse-problem half is not yet finished.

  2. II Early

    What can be known together?

    Distributed, partial observers

    One good camera is a single point of failure and a privacy problem. Many cheap, unreliable, intermittent observers might be the better instrument — but only if they can agree on identity when they have never synchronised, and on what to do when their views conflict. A shared spatial state has to survive disagreement, latency and loss.

    Formally open. Hydra-Bench is a first step toward synchronised multi-modal evidence.

  3. III Open

    When should intelligence change the world?

    Authority, and the right to decline

    A wrong action has a physical cost that a wrong prediction does not. Reasoning should be able to propose, revise and defer, while the decision to actually move something stays inside explicit bounds. I want to know when better prediction does not change the decision at all, and when a system should hand control back to a person instead of acting.

    Driven by the ageing-and-driving work, where the cost of a false alarm and the cost of a miss are not the same number.

These questions were not chosen and then fitted to the work. Each one is where a project left me — the leaf-wetness line forced observability and persistence, the driving line forced meaning and authority, and the two together forced the question of who decides.

What I am looking for

I am not optimising for a job title. I want a setting where the sensing, the modelling and the decision are argued about by people who disagree — and where the work has to survive contact with somewhere real.

08 The ambition

What I actually
want to build

Everything above is a piece of the problem taken on its own. This is the thing I actually want to exist at the end of it.

  1. 01 Many partial observers — a camera, a radio return, a car, a sensor on a wall — each seeing a fragment, each failing in its own way.
  2. 02 One shared spatial state that survives their disagreement, their latency and their absence — and knows what it does not know.
  3. 03 A single committed action taken from that state, with the right to decline to act at all.

I want to build the runtime that makes this possible: a shared physical-world understanding that many bounded observers can contribute to and rely on, where the state carries its own uncertainty, and where the decision to change the world is separable from the reasoning that proposed it. Not a better model — a shared substrate that models can be composited onto, the way sensors are composited onto a place.

What does not exist yet

None of this exists yet in the form I have just described. What exists is the sensing and representation half of step 01 and 02, published and citable; step 03 exists as an argument rather than as a system. The work between here and there is the part I am actually asking a lab for.

06 Still arguing

Questions I do not
have an answer to

These are questions, not conclusions. They are the part of my work that has no paper attached yet, and the part I expect to still be arguing about in five years.

  1. Q1 Open question

    What does a model owe the space it models?

    If understanding a space means maintaining a belief about it, then it also means representing what the belief is missing. How should a learned world state carry its own ignorance — and how would we ever measure it?

    • Occlusion as first-class state
    • Calibrated uncertainty over claims, not over pixels
    • Evaluation where the ground truth is, by construction, unobservable
  2. Q2 Early work

    Sensor networks as the nervous system of built space

    One good camera is a single point of failure and a privacy problem. Many cheap, unreliable, intermittent observers might be better. What does an architecture look like when perception is distributed, asynchronous, and mostly missing?

    • Redundancy across modalities, locations and time
    • Identity when the same thing is seen by sensors that never agree
    • Energy, maintenance and the real cost of watching a place
  3. Q3 Speculative

    Physical intelligence without a body

    A car, a wheeled robot and a radar bolted to a wall all have to understand space. Is there an abstraction above embodiment that is not merely a policy, and what would separate a model of a place from a habit of moving through it?

    • What actually transfers between embodiments
    • State that is useful without being actionable
    • Grounding as a design constraint rather than a feature
  4. Q4 Open question

    When should a system ask instead of act?

    In safety-critical space a false alarm has a social cost and a missed one has a physical one. How should an agent price that asymmetry, and what does principled abstention look like when the correct behaviour is to hand control back to a human?

    • Abstention and deferral in driving and care
    • Risk that is legible to the person being put at risk
    • Evaluation that rewards early gentle intervention

Standing notes

  • Sensor fusion is usually framed as improving accuracy. I am more interested in the opposite: what a system has to give up in order to stay alive when half of its inputs are gone.
  • The hard part of a world model is not the future. It is remembering what it could not see.
  • A benchmark reporting average accuracy says nothing about the moment a pedestrian steps out of occlusion.