Again AI Claude assisted πŸ™‚

@cfrank This is a really generous reply, and it corrects me on two things I got wrong, so let me start there.

Your objective was already utility-maximizing for each faction, not CW-displacement β€” so the question I made the most noise about was one you'd already answered. And your generator wasn't impartial culture, it was stress profiles. Both of my sharpest points were aimed at problems you didn't have. Apologies for assuming.

I spent some time with your plots, and I think they make your own case more precisely than the summary does β€” including in one place where they cut against your original conclusion.

The crossover only happens at the PSRO endpoint. Reading off your panels (so Β±1 point): at 0% strategic the Smith methods are at 100% CW rate and 0.992 normalized utility against IRV's ~59% and 0.944. At 25/50/75/90% they're at ~90/81/74/70% and 0.987β†’0.955, against IRV flat at ~61–64% and ~0.937–0.944. So on both the hit rate and the utility measure, Smith methods dominate IRV at every level of strategic participation you tested short of the endpoint β€” including 90%, where nine voters in ten are strategic. That's a much stronger statement than "the sincere advantage can remain large," and it's your data, not mine.

Your "outside sincere Smith rate" panel is the cleanest thing in the set. IRV sits flat at roughly 36–40% at every level, including 0%. It elects outside the sincere Smith set about 40% of the time with nobody lying at all. That flat line is the whole reason the convergence appears: strategy has very little left to take from IRV, because IRV has already spent it. The Smith methods start at 0% and are driven up to 25–30% by strategy; IRV starts at 40% and stays. Those are different failures wearing the same number.

Two smaller things. First, I think there may be an off-by-one between your prose and your plot: you quote Ranked Pairs at 97/90/81/74 for 25/50/75/90%, but the plot reads roughly 90/81.5/73.5/70 at those positions β€” your sequence looks shifted one notch. Doesn't change the shape, but at 90% strategic it's ~70%, not 74%.

Second, and I think this matters for the conclusion: at the PSRO endpoint your best method is IRV + fresh runoff, not IRV. It leads on all three panels β€” ~69% CW rate, 0.964 utility, and the lowest outside-sincere-Smith rate of any method at that endpoint (~31%, while plain IRV is at ~39.5% and Ranked Pairs is at ~56%). So even taking the adaptive regime entirely at face value, the finding isn't "IRV is robust to sophisticated strategy." It's "a second round on fresh ballots is robust to sophisticated strategy" β€” a claim about two-round structure rather than about Smith compliance, and one that would apply just as well to score-plus-runoff designs. I'd be curious whether that holds up as you vary the runoff's ballot.

On your mechanism β€” the CW being pushed out of the Smith set before completion runs. That's a sharp claim and I hadn't separated it out, so I measured it. Single coordinated bloc, burial, rational (it only submits if it beats voting honestly), every challenger tried, swept over 3/5/7/9 candidates at your 71 voters. Ejection is the minority regime everywhere: 19–36% of successful burials remove the CW from the reported Smith set, and the other two-thirds to four-fifths leave the CW sitting inside it, where the completion rule still decides. Interestingly the ejected share barely moves with field size (23%β†’26% in 1-D from 3 to 9 candidates) even as raw displacement climbs from 20% to 88%. A wider field makes burial much easier without much changing where it lands.

The obvious limit: that's one bloc doing a plain bury-to-last, not several factions best-responding over rank and score offsets. Ejection is clearly something a stronger search buys. Which I think locates our remaining disagreement precisely and answerably: how much of the ejection rate is purchased by adaptive multi-faction optimization, over and above single-bloc burial? If PSRO is ejecting at 60–70% where a naive bloc ejects at 25%, that's a real and quotable finding about what adaptivity does, and it would be worth a paper on its own. Code's here if it's useful: https://masiarek.github.io/star-voting-library/07_Concepts/topics/compliance_vs_strategic_preservation.html

On the AI point β€” I think this is the most interesting thing in your post and I don't want to wave it away, because the direction is right. Three things give me pause about how far it goes:

Computation isn't the binding constraint. In my runs a successful burial needed 34–40% of the electorate to rank someone they genuinely like dead last. AI can tell you that's the optimal play; it can't get 40% of voters to cast it on trust. That cost is social, and it's the one that doesn't fall with compute.

The information required is about opponents, not preferences. PSRO converges because it iterates against a live opponent mixture, observing what the other factions actually do. A real electorate votes once, without seeing anyone's final policy. Better tools don't close that gap β€” they arguably widen the variance, since everyone is now optimizing against a guess about everyone else, and a burial aimed at the wrong equilibrium is exactly the one that backfires.

Cheaper attacks make cheaper detection, and that cuts toward pairwise methods. A successful burial's signature is a cycle appearing in a race whose pre-election pairwise polling showed a clean head-to-head winner. The thing detection needs is a published pairwise matrix β€” which is precisely what Condorcet methods emit as a byproduct and what IRV doesn't. If the threat model is "strategic optimization gets cheap," I'd want the method that publishes the most auditable structure, not the least.

None of which touches your core point, which I think is right and which I've now written up: formal compliance is a property of the cast ballots, "the sincere winner wins" is a property of the electorate plus its behaviour, and treating the first as a guarantee of the second is sloppy. That's a correction advocates on my side of this should absorb rather than argue with. I just don't think it demotes Smith compliance β€” your own 0–90% panels are about as strong an argument for it as I've seen.