@masiarek thank you, I appreciate your feedback.
First, to clarify the simulation: ChatGPT used lightweight PSRO-like strategic agents representing coordinated factions. Each faction’s objective was to maximize its members’ mean underlying utility for the elected candidate, not to defeat the sincere Condorcet winner. An evolutionary best-response procedure searched over strategically available ballot policies against the current mixture of opposing strategies. This was done in spare time with ChatGPT, so I don’t have the code base yet, but can share once I dig it out.
The three largest sincere-favorite factions were strategic agents. Their policies controlled strategic participation/sincerity and candidate-specific rank and score offsets. I used 9 candidates and 71 voters per election.
Also, the benchmark was not impartial culture. I was using synthetic stress profiles including center-squeeze, polarized, clone-heavy, and ring-type electorates. So your broader point about generator dependence still applies, but the ~70% IRV result wasn’t coming from IC.
So, in that respect, a manipulation that displaced the sincere Condorcet winner but produced a worse outcome for the manipulating faction would count against that strategy. That said, model dependence is definitely a serious limitation.
Second, I did test a fairly broad range of completion rules for Smith-compliant methods rather than treating “Smith” as a single method. These included Ranked Pairs, B2R, IRV, Approval, Score, Benham, and several other standard or experimental completions. I also tested variants using fresh second-round ballots, including cases where the second-round completion was itself Smith-restricted. I may certainly have missed some more obscure possibilities, but the convergence did not seem specific to one particular Smith completion.
The main failure mode was often burial or burial-like pairwise distortion: the sincere Condorcet winner had already been pushed out of the Smith set before the completion rule was applied, so at that point changing the completion rule could not rescue it, although different completion rules might change some strategic incentives and definitely impact results once the sincere CW makes it into the Smith set.
I agree with your assessment in an important respect. These simulations assume unusually capable, coordinated strategic factions. They are basically searches for strategies that sophisticated actors could discover, rather than models or predictions of what ordinary voters would actually do, discover, or successfully coordinate.
There is also a normative question here. A system can perform relatively well once voters behave ruthlessly and strategically while performing worse under sincerity, but it is not obvious that this is desirable, and it is unlikely to be descriptively realistic. Obviously, effective strategy has informational, computational, coordination, and cognitive costs that many voters will not pay—or may not be able to pay equally.
I did some follow-up experiments after reading your response where I restricted behavior to be more “human-like” in the loose sense that strategic agents used simpler heuristic strategies, and I varied the fraction of strategic voters. The results were quite different: conditioning on elections with a unique sincere Condorcet winner, Ranked Pairs elected that winner about 97%, 90%, 81%, and 74% of the time when 25%, 50%, 75%, and 90% of voters, respectively, were strategic. IRV remained around 63–65% across those conditions.
You can see these plotted here.
So I think the distinction you’re drawing is important, and I largely agree. The earlier convergence between IRV and Condorcet methods seems to be a result about sufficiently powerful adaptive strategic behavior, not about strategic voting in general, or in human circumstances. Under simpler and more heterogeneous strategic behavior, the sincere advantage of Condorcet methods can remain very large.
However, I don’t think the highly adaptive case is therefore unimportant. As AI tools become increasingly capable and accessible, the informational and computational costs of identifying sophisticated election strategies may decline substantially. So even if the adaptive simulations are poor descriptive models of present-day individual voter behavior, they may still be useful as stress tests of how a voting method behaves when strategic optimization becomes easier.