3D Intention Grounding · Functional Reasoning · 3D Vision-Language Understanding
CoCR bridges abstract human intentions and 3D target objects through an explicit intent → function → object reasoning chain.
Decomposes abstract intentions into step-wise latent needs and concrete functional requirements.
Builds explicit functional dependency paths linking the intent, functional nodes, and candidate objects.
Aligns structured functional reasoning with 3D visual evidence and iteratively refines the dependency graph.
From abstract intent parsing to function-aware 3D object selection.
A fine-tuned T5-small model generates a step-wise reasoning sequence. Each reasoning step is mapped into a unified functional space containing 40 functional requirement nodes.
Functional nodes are linked to candidate objects using normalized compatibility, yielding an explicit intent–function–object dependency graph with causal pruning.
Cross-attention injects functional cues into 3D object features, while dependency-weighted consistency supports mutual verification and iterative graph updates.
Results on the Intent3D benchmark.
| Method | Val Top1@.25 | Val Top1@.5 | Val AP@.25 | Val AP@.5 | Test Top1@.25 | Test Top1@.5 | Test AP@.25 | Test AP@.5 |
|---|---|---|---|---|---|---|---|---|
| BUTD-DETR | 47.12 | 24.56 | 31.05 | 13.05 | 47.86 | 25.74 | 31.41 | 13.46 |
| 3D-ViSTA | 42.76 | 30.37 | 36.10 | 19.93 | 43.88 | 31.44 | 37.29 | 22.00 |
| IntentNet | 58.34 | 40.83 | 41.90 | 25.36 | 58.92 | 42.28 | 44.01 | 27.60 |
| CoCR (Ours) | 60.71 | 43.24 | 42.65 | 27.33 | 61.37 | 45.91 | 45.73 | 30.28 |
Animated demonstrations, successful localization, and failure modes.
Reasoning depth and graph update iterations.