NeurIPS 2026

Detect What You NeedChain-of-Causal Reasoning for 3D Intent Grounding

Zihao Zhang Aming Wu Yang Li Yahong Han

3D Intention Grounding · Functional Reasoning · 3D Vision-Language Understanding

Paper Method Visualizations

CoCR bridges abstract human intentions and 3D target objects through an explicit intent → function → object reasoning chain.

Motivation. Direct semantic matching can introduce a logical leap. CoCR explicitly infers functional requirements before grounding candidate 3D objects.
Overview

Abstract

3D Intention Grounding aims to localize target objects in 3D scenes from natural-language intentions. Unlike conventional visual grounding, the language is often abstract and non-descriptive, requiring models to infer latent functional requirements before matching them with object-level 3D evidence. We introduce Chain-of-Causal Reasoning (CoCR), a causality-inspired functional dependency reasoning framework that explicitly connects intentions and candidate objects through intermediate functional requirements. CoCR progressively decomposes complex intents, constructs an interpretable intent–function–object dependency graph, and aligns function-aware representations with geometric-semantic point-cloud features. Experiments on 3D Intention Grounding and 3D Visual Grounding demonstrate improved intent-aware object localization.
1

Intent Causal Parsing

Decomposes abstract intentions into step-wise latent needs and concrete functional requirements.

2

Intent Causal Grounding

Builds explicit functional dependency paths linking the intent, functional nodes, and candidate objects.

3

Causal–Visual Alignment

Aligns structured functional reasoning with 3D visual evidence and iteratively refines the dependency graph.

Method

Chain-of-Causal Reasoning

From abstract intent parsing to function-aware 3D object selection.

Overall architecture. CoCR jointly processes the intent text and 3D point cloud, constructs a functional dependency graph, performs causal–visual feature alignment and pruning, and localizes the target object through the refined graph.

Intent Causal Parsing (ICP)

A fine-tuned T5-small model generates a step-wise reasoning sequence. Each reasoning step is mapped into a unified functional space containing 40 functional requirement nodes.

Intent Causal Grounding (ICG)

Functional nodes are linked to candidate objects using normalized compatibility, yielding an explicit intent–function–object dependency graph with causal pruning.

Causal–Visual Feature Alignment (CVFA)

Cross-attention injects functional cues into 3D object features, while dependency-weighted consistency supports mutual verification and iterative graph updates.

Experiments

Quantitative Results

Results on the Intent3D benchmark.

3D Intention Grounding on Intent3D
MethodVal Top1@.25Val Top1@.5Val AP@.25Val AP@.5Test Top1@.25Test Top1@.5Test AP@.25Test AP@.5
BUTD-DETR47.1224.5631.0513.0547.8625.7431.4113.46
3D-ViSTA42.7630.3736.1019.9343.8831.4437.2922.00
IntentNet58.3440.8341.9025.3658.9242.2844.0127.60
CoCR (Ours)60.7143.2442.6527.3361.3745.9145.7330.28
Qualitative Study

Visualizations

Animated demonstrations, successful localization, and failure modes.

Animated examples. Target localization across four scenes and natural-language intentions.
Comparison with IntentNet. CoCR localizes all target sinks in the first example and correctly distinguishes the TV stand from functionally similar objects in the second.
Successful localization across diverse scenes and objects.
Failure cases show how visually or functionally similar objects can still lead to incorrect localization.
Analysis

Hyperparameter Analysis

Reasoning depth and graph update iterations.

Four reasoning steps and two graph-update iterations yield the strongest performance in these analyses.