Approximate semantic-geometric correspondences generated by SemGeo-Gen between two different cars, then two different airplanes.

Abstract

We introduce SemGeo-Gen, an unsupervised method for generating approximate cross-instance semantic-geometric supervision from weakly aligned 3D object collections. Our goal is not to recover exact point-level correspondences, which are often ambiguous across different instances and would require expensive, impractical manual supervision at scale. Instead, we automatically produce approximate correspondences that are semantically meaningful, geometrically consistent, and diverse enough to train modern correspondence models. Given only coarse category-level rotation alignment, SemGeo-Gen lifts multi-view DINOv2 features to 3D, aligns object instances using continuous piecewise-affine registration, and prunes candidate matches using semantic consistency. The resulting 3D correspondences can be projected into rendered views, yielding scalable 3D–3D, 3D–2D, and 2D–2D supervision without manual annotation. We validate the generated correspondences against sparse human annotations from KeypointNet, obtaining 93% PCK@0.10 and 97% PCK@0.15. More importantly, we show that using the generated correspondences for synthetic pretraining improves recent state-of-the-art semantic correspondence models on SPair-71k after fine-tuning on real data, with gains of up to 5 PCK@0.10 points. Together, these results indicate that approximate semantic-geometric supervision generated from 3D assets can improve real-image correspondence learning and serve as a scalable alternative to costly manual annotation.

0
manual labels (no point, part, or correspondence annotation)
93%
PCK@0.10 agreement with human KeypointNet annotations
+4.4 / +2.9
mean PCK@0.10 on SPair-71k Rigid-9 for GeoAware-SC / MARCO

One Generator, Three Modalities

Matches are established in 3D and projected into rendered views, giving 3D–3D, 3D–2D, and 2D–2D supervision.

2D-2D correspondences between rendered views of two airplanes
2D–2D
3D-2D correspondences between a 3D airplane and a rendered view of another
3D–2D
3D-3D correspondences between the surfaces of two airplanes
3D–3D

Approximate cross-instance correspondences between two airplane instances. Colors in 3D visualize the lifted DINOv2 features reduced to three dimensions.

Method

Input: only rotation-aligned object pairs, with no point, part, or correspondence labels.

SemGeo-Gen pipeline: CPA registration of the source onto the target, establishing correspondences, and semantic correspondence pruning

1Lift semantic features to 3D

Each object is rendered from calibrated virtual cameras, dense DINOv2 features are extracted in image space and lifted back to the visible 3D points, and multi-view observations are aggregated into one descriptor per point. Semantic compatibility is the cosine similarity of the descriptors after a joint 64-D PCA:

\[ s_{\mathrm{sem}}(x_i,x_j)=\frac{z_i^\top z_j}{\|z_i\|_2\,\|z_j\|_2} > \tau_{\mathrm{sem}}, \qquad \tau_{\mathrm{sem}}=0.85. \]

2Semantic-coverage pair filtering optional

A pair is kept only if, in both directions, a large enough fraction of points has a semantically compatible point in the other object:

\[ \rho_{\mathrm{sem}}(\mathcal{S}\!\rightarrow\!\mathcal{T})=\frac{1}{|\mathcal{X}_S|}\sum_{x_i^S\in\mathcal{X}_S}\mathbf{1}\Big[\max_{x_j^T\in\mathcal{X}_T} s_{\mathrm{sem}}(x_i^S,x_j^T)>\tau_{\mathrm{sem}}\Big] \;\ge\; \gamma=0.6 . \]

3Continuous piecewise-affine (CPA) registration

The source is deformed onto the target by a control grid over the normalized 3D domain, decomposed into tetrahedra by Delaunay triangulation. Each vertex has a learnable displacement \(u_v\), and points move by barycentric interpolation, so the deformation is affine inside each tetrahedron and continuous across cells. It is fit by gradient descent on a Chamfer loss with smoothness regularization:

\[ \min_{T_{\mathrm{CPA}}}\; \mathcal{L}_{\mathrm{Chamfer}}\big(T_{\mathrm{CPA}}(\mathcal{X}_S),\mathcal{X}_T\big) + \lambda_{\mathrm{smooth}}\,\frac{1}{|\mathcal{E}|}\sum_{(v_a,v_b)\in\mathcal{E}}\|u_{v_a}-u_{v_b}\|_2^2, \qquad \lambda_{\mathrm{smooth}}=0.3, \] \[ \mathcal{L}_{\mathrm{Chamfer}}(\widehat{\mathcal{X}}_S,\mathcal{X}_T)=\frac{1}{|\widehat{\mathcal{X}}_S|}\sum_{\hat{x}\in\widehat{\mathcal{X}}_S}\min_{x\in\mathcal{X}_T}\|\hat{x}-x\|_2^2+\frac{1}{|\mathcal{X}_T|}\sum_{x\in\mathcal{X}_T}\min_{\hat{x}\in\widehat{\mathcal{X}}_S}\|x-\hat{x}\|_2^2, \]

where \(\mathcal{E}\) are the edges between neighboring vertices of the tetrahedral control grid.

4Match, then prune

Each deformed source point \(\hat{x}_i^S=T_{\mathrm{CPA}}(x_i^S)\) is assigned its nearest target point. The stored match links the original source point to that target point, and is kept only if it is also semantically compatible:

\[ a(i)=\arg\min_{j}\big\|\hat{x}_i^S-x_j^T\big\|_2, \qquad \mathcal{C}_{\mathcal{S},\mathcal{T}}=\Big\{\big(x_i^S,x_{a(i)}^T\big)\;:\;s_{\mathrm{sem}}\big(x_i^S,x_{a(i)}^T\big)>\tau_{\mathrm{sem}}\Big\}. \]

5Project to 2D, scale combinatorially

Rendering with calibrated cameras \(\pi_{c_S},\pi_{c_T}\) turns each 3D match into 3D–2D (one endpoint projected) or 2D–2D (both projected) supervision:

\[ \big(u_i^S,u_j^T\big)=\big(\pi_{c_S}(x_i^S),\,\pi_{c_T}(x_j^T)\big). \]

A category with \(N\) objects yields up to \(\binom{N}{2}\) cross-instance pairs, and each pair can be rendered from many views.

Pretraining Improves State-of-the-Art Models

Pretrain on 12K generated 2D pairs per category from Objaverse-OA, then fine-tune on real SPair-71k. No architectural changes.

MethodTraining regimeaerobikeboatbuscarchairmbiketraintvMean
Unsupervised / weakly supervised methods
SDSPair-71k62.852.731.239.135.632.051.061.852.946.57
DINOv2SPair-71k73.460.243.246.745.133.460.754.223.948.98
DINOv2+SDSPair-71k73.861.040.247.444.141.561.763.552.453.96
SphericalMapsSPair-71k76.260.146.574.968.045.169.173.958.163.54
Supervised methods
SCorrSanSPair-71k57.140.338.157.847.125.245.377.769.750.92
CATS++SPair-71k60.646.941.664.950.429.250.980.974.955.59
DHFSPair-71k74.061.040.770.074.438.566.687.460.363.66
DINO+SD (S)SPair-71k84.767.564.585.782.057.075.993.670.575.71
Jamais VuSPair-71k83.670.164.191.483.261.478.495.580.178.80
Recent supervised state-of-the-art methods under different supervision settings
GeoAware-SCSPair-71k83.4271.0766.6890.9386.1062.1879.5495.3379.0079.36
SPair-71k + Ours86.2874.6068.3393.1284.7368.9479.7395.3180.7381.31
Ours P.T. → SPair-71k89.3876.1171.1593.8086.7473.8082.3395.6784.5983.73
MARCOSPair-71k92.1676.7676.5393.1788.5782.4982.0393.6291.7286.39
SPair-71k + Ours92.2677.4973.7492.5984.8682.7179.7893.2289.7285.23
Ours P.T. → SPair-71k95.5980.6679.3594.5991.8991.1682.3094.6993.4989.28

Category-wise PCK@0.10 ↑ on Rigid-9: the rigid SPair-71k categories with well-defined semantic-geometric structure (excluding the rotationally symmetric bottle and potted plant). P.T. = pretraining.

Rigid-only pretraining also helps non-rigid categories

MethodTrainingRigid-9Non-rigid-7Symmetric-2All-18
GeoAware-SCSPair-71k80.4886.9472.3582.09
DenseMatcher P.T.80.3786.5471.9781.83
Ours P.T.84.20 +3.7288.23 +1.2972.58 +0.2384.48 +2.39
MARCOSPair-71k87.3290.7974.9987.30
Ours P.T.89.61 +2.2992.06 +1.2775.38 +0.3988.98 +1.68

Pretraining on synthetic rigid categories, then fine-tuning on all 18 SPair-71k categories. PCK@0.10 ↑.

Pseudo-label quality matters

Pretraining labelsRigid-9 mean
None (SPair-71k only)79.36
DenseMatcher80.13 +0.77
SemGeo-Gen (ours)83.73 +4.37

GeoAware-SC, same data-generation and pretrain-then-fine-tune protocol. PCK@0.10 ↑.

Qualitative Results on SPair-71k

Each image shows a source–target pair with predicted correspondences. Green: correct; red: incorrect.

Semantic Coverage & Pruning

Geometry alone matches nearby regions even when they have no semantic counterpart; lifted DINOv2 features remove those matches.

Rendering
Joint DINOv2 PCA
Semantic coverage
Roofless car: rendering, joint PCA features, coverage mask with red seats White car: rendering, joint PCA features, all-green coverage mask Roofless car from a second view White car from a second view

Green points have at least one point in the paired object with cosine similarity above 0.85; red points have none (e.g., the exposed seats) and receive no correspondences.

Final correspondences after pruning

Sampled final correspondences between the two cars, view 1 Sampled final correspondences between the two cars, view 2

Sampled correspondences after semantic pruning, from two views of the same pair. The exposed seats and steering wheel, and the roof of the white car, receive no correspondences even though registration deforms them toward other parts.

Agreement with Human Annotations

Generated 3D correspondences vs. human-annotated KeypointNet keypoints, used for evaluation only. Three subsets of 100 models per category.

Category@0.05@0.10@0.15@0.20
Airplane0.790.940.980.99
Bathtub0.700.880.960.99
Bed0.680.910.960.99
Bottle0.830.960.980.99
Cap0.800.971.001.00
Car0.780.950.991.00
Chair0.770.960.990.99
Guitar0.810.960.980.99
Helmet0.550.840.940.98
Category@0.05@0.10@0.15@0.20
Knife0.830.930.960.98
Laptop0.890.991.001.00
Motorcycle0.750.930.980.99
Mug0.730.960.991.00
Skateboard0.860.970.991.00
Table0.760.960.991.00
Vessel0.450.710.870.94
Mean0.750.930.970.99
Median0.780.960.980.99

PCK at thresholds 0.05–0.20 of the target bounding-box diagonal ↑.

Comparison with Back to 3D

Method@0.05@0.10@0.15@0.20
Back to 3D (1-shot)0.4380.6320.7250.777
Back to 3D (3-shot)0.4360.6410.7400.799
SemGeo-Gen (zero-shot)0.6480.8650.9450.977

Mean PCK on a subset of 100 source–target pairs per category (1,600 pairs).

Semantic filtering ablation

Setting@0.05@0.10@0.15@0.20
Registration only (no filtering)0.620.840.930.97
+ correspondence pruning0.730.910.970.99
+ pair filtering (full)0.750.930.970.99

Mean PCK on KeypointNet. Pair-level numbers are computed on the retained pairs.

Registration model ablation

ModelDOFTime (s/pair)@0.01@0.05@0.10@0.15@0.20
Affine121.150.2210.6860.8980.9600.982
TPS (K=216)6603.710.2260.7390.9190.9660.983
CPA (ours)6481.130.2360.7290.9220.9710.986

Mean PCK across the 16 KeypointNet categories; registration time on a single A6000 GPU. CPA matches TPS accuracy while running about 3× faster.

Beyond Existing Datasets

Correspondences on Objaverse-OA categories not covered by SPair-71k, KeypointNet, or AP-10K.

Correspondences between two baby buggies Correspondences between two backpacks Correspondences between two barbell benches

Citation

If you find this work useful, please cite:

@inproceedings{Amoyal:NeurIPS:2026:SemGeoGen,
      title={{SemGeo-Gen: Unsupervised Generation of Approximate Cross-Instance Semantic-Geometric Correspondences}},
      author={Amoyal, Roy and Ifergane, Shira and Freifeld, Oren},
      year={2026},
      booktitle={NeurIPS},
}