Skip to content

Commit c889af3

Browse files
Update performance note
1 parent 47cf0f2 commit c889af3

1 file changed

Lines changed: 43 additions & 12 deletions

File tree

development/graph/lifted_multicut/PERFORMANCE_NOTES.md

Lines changed: 43 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -16,23 +16,54 @@ is faster.
1616

1717
| Problem | Solver | bic energy | nifty energy | Δenergy | bic runtime | nifty runtime | runtime ratio |
1818
|---|---|---|---|---|---|---|---|
19-
| 2D | greedy | -1575.04 | -1575.04 | 0.00 | 0.70 ms | 1.50 ms | 2.15× faster |
20-
| 2D | KL (10 outer) | -1575.21 | -1575.21 | 0.00 | 4.14 ms | 4.77 ms | 1.15× faster |
21-
| 2D | fusion-move | -1575.43 | -1575.43 | 0.00 | 8.31 ms | 12.3 ms | 1.48× faster |
22-
| 3D | greedy | -15891.4 | -15891.0 | −0.35 | 6.40 ms | 15.9 ms | 2.48× faster |
23-
| 3D | KL (10 outer) | -15921.0 | -15921.1 | +0.07 | 78.1 ms | 103 ms | 1.32× faster |
24-
| 3D | fusion-move | -15915.1 | -15915.1 | 0.00 | 128 ms | 196 ms | 1.53× faster |
25-
| grid | greedy | -690 014 | -690 050 | +35.9 | 16.4 s | 20.6 s | 1.25× faster |
26-
| grid | fusion-move | -690 271 | -690 356 | +84.7 | 51.8 s | 60.2 s | 1.16× faster |
19+
| 2D | greedy | -1575.04 | -1575.04 | 0.00 | 0.70 ms | 1.50 ms | 2.15× faster |
20+
| 2D | KL (10 outer) | -1575.21 | -1575.21 | 0.00 | 4.14 ms | 4.77 ms | 1.15× faster |
21+
| 2D | fusion-move | -1575.43 | -1575.43 | 0.00 | 8.31 ms | 12.3 ms | 1.48× faster |
22+
| 3D | greedy | -15891.4 | -15891.0 | −0.35 | 6.40 ms | 15.9 ms | 2.48× faster |
23+
| 3D | KL (10 outer) | -15921.0 | -15921.1 | +0.07 | 78.1 ms | 103 ms | 1.32× faster |
24+
| 3D | fusion-move | -15915.1 | -15915.1 | 0.00 | 128 ms | 196 ms | 1.53× faster |
25+
| grid | greedy | -690 014 | -690 050 | +35.9 | 16.4 s | 20.6 s | 1.25× faster |
26+
| grid | KL (10 outer) | -690 544 | -690 599 | +54.5 | 442 s | 6 514 s | **14.7× faster** |
27+
| grid | fusion-move | -690 271 | -690 356 | +84.7 | 51.8 s | 60.2 s | 1.16× faster |
2728

2829
Δenergy = bic − nifty; negative means bic is better, positive means
2930
nifty is better. Energies are exact matches on 2D/3D fusion-move and
3031
within 0.05 % on the rest; bic is faster than nifty on every row.
3132

32-
KL on the 262 k-node grid is omitted from the matrix — it is correct
33-
but takes several minutes (heavy chain-init work scales with cluster
34-
count). See the fusion-move post-script below for the grid-specific
35-
behavior.
33+
### Note on the grid-KL row
34+
35+
The 14.7× speed gap and the 54.5-unit energy gap on this single row are
36+
much larger than on every other row (the 2D / 3D KL ratios are 1.15–1.32×
37+
with essentially exact energies). That points to nifty's KL doing more
38+
work per outer iter on grid — not a single algorithmic difference, but
39+
the most plausible drivers, given that no investigation has been done:
40+
41+
- **Inner-loop epsilon.** Our two-cut driver runs with one `epsilon`
42+
parameter, defaulted to `1e-6` from
43+
`KernighanLinSolver` (header file `kernighan_lin.hxx`). nifty's outer
44+
KL uses `epsilon=1e-7` and its inner `TwoCut` uses
45+
`epsilon=1e-9`. A tighter epsilon admits smaller-gain moves, which
46+
means longer chains per pair and more outer iters before the
47+
no-improvement break. On a 262 k-node problem with millions of
48+
candidate moves, the extra accepted moves multiply quickly.
49+
- **Inner iteration cap.** Both implementations effectively run the
50+
two-cut chain until the heap empties; in nifty this is governed by
51+
`numberOfInnerIterations` defaulting to
52+
`std::numeric_limits<std::size_t>::max()`. Same effective contract
53+
as bic, but if a different value is wired through the factory the
54+
gap would shift.
55+
- **Per-pair adjacency cost.** Our `chain_gain_init` builds a filtered
56+
per-node adjacency cache and the inner loop iterates only in-pair
57+
neighbours. nifty walks the full lifted adjacency on every move and
58+
re-classifies in/out-of-pair via label comparisons. At the scale of
59+
grid (long-range lifted edges, dense pair adjacency), the constant
60+
factor here adds up to multiple-× per chain — independent of
61+
whether the algorithm visits the same number of moves.
62+
63+
Two of the three (epsilon, adjacency walking) push energy and runtime
64+
in the same direction: nifty does strictly more inner work and lands at
65+
a slightly better energy. Closing the 54.5-unit gap on bic would mean
66+
tightening `epsilon` and accepting more runtime — not pursued here.
3667

3768
## What's done
3869

0 commit comments

Comments
 (0)