The first time I heard of L0, L1, L2 strategy was from David Sklansky's Theory of Poker book. Many discussions of leveling occured on twoplustwo's poker forums.
Thanks! Part of me wishes that I picked up a more lucrative hobby in my twenties like online poker instead of blogging, RPS bots, and random board games. But realistically I know I'd either lose money, or worse, grind away all my spare time to make in expectation 5 dollars an hour lol.
most people are in a card game for other reasons that profitable play. also important to understand that reinvestment of capital does not require genius level thinking to embrace. this means there are a lot of rich dummies that are happy to dump money. however, it does no good to beat the fish and lose to tougher pros. thus gto solutions to break even vs the pros. all this requires a bankroll. thats typically the hard part. and realiatic self evaluation. lol. the real hard part!
I'm curious what you think they are! I think the most interesting thing about RPS is that the game is simple enough that you can study and understand the adversarial dynamics in a fairly pure form.
I suspect some lessons from it generalize well to other adversarial dynamics including but not limited to other games, gambling, financial markets, and some aspects of war. In particular the tension between trying to maximize your rewards from exploiting weaker opponents while avoiding exploitation yourself seems like something that might generalize well.
But details matter a lot so application is far from straightforward.
Digits are not uniform mod 3 since 3 does not divide 10, so your human “pure random” strategy is rock-biased. Can fix by skipping to the next digit on 9s.
"Obviously a base 10 rendition of pi has some biases mod 3. Fortunately “0” does not show up in pi until the 32nd digit, long after most people stop playing."
Pseudorandomness can be easy for machines, but LLMs are very bad at it (even at high temperature). Try asking a 2025 LLM to flip a coin, it's almost always heads.
I just did a quick spot check on the three Anthropic 4.5 models; across six games I got five scissors and one rock.
Eg if you ask them to try pretty hard to win, or if they read this article (or a summary of it), will do they better?
If not, then we have an odd situation where a fairly simple specialized bot (eg Henny, or string-finder + Henny) can beat a SOTA general LLM! Surprising if true.
I'm going to test this when I get a chance. For now I'll register a prediction: for any given current frontier LLM and commonly-used sampler with a fixed context window using one of the approaches you described, the most common pick will still occur at least ⅔ of the time.
Interested in a complexity-weighted RPS contest! Fantastic primer, and well written. Thanks for sharing :)
You're very welcome! :D
The first time I heard of L0, L1, L2 strategy was from David Sklansky's Theory of Poker book. Many discussions of leveling occured on twoplustwo's poker forums.
Have fun storming the castle!
Thanks! Part of me wishes that I picked up a more lucrative hobby in my twenties like online poker instead of blogging, RPS bots, and random board games. But realistically I know I'd either lose money, or worse, grind away all my spare time to make in expectation 5 dollars an hour lol.
most people are in a card game for other reasons that profitable play. also important to understand that reinvestment of capital does not require genius level thinking to embrace. this means there are a lot of rich dummies that are happy to dump money. however, it does no good to beat the fish and lose to tougher pros. thus gto solutions to break even vs the pros. all this requires a bankroll. thats typically the hard part. and realiatic self evaluation. lol. the real hard part!
interested in discussion of real-world applications (i think there are a bunch)
I'm curious what you think they are! I think the most interesting thing about RPS is that the game is simple enough that you can study and understand the adversarial dynamics in a fairly pure form.
I suspect some lessons from it generalize well to other adversarial dynamics including but not limited to other games, gambling, financial markets, and some aspects of war. In particular the tension between trying to maximize your rewards from exploiting weaker opponents while avoiding exploitation yourself seems like something that might generalize well.
But details matter a lot so application is far from straightforward.
I'm curious if you have other ideas/suggestions!
(Other commenters welcome to chime in, of course!)
Digits are not uniform mod 3 since 3 does not divide 10, so your human “pure random” strategy is rock-biased. Can fix by skipping to the next digit on 9s.
This is addressed in the footnote!
"Obviously a base 10 rendition of pi has some biases mod 3. Fortunately “0” does not show up in pi until the 32nd digit, long after most people stop playing."
Pseudorandomness can be easy for machines, but LLMs are very bad at it (even at high temperature). Try asking a 2025 LLM to flip a coin, it's almost always heads.
I just did a quick spot check on the three Anthropic 4.5 models; across six games I got five scissors and one rock.
Interesting! I wonder if it's fixable with light prompting.
Eg if you ask them to try pretty hard to win, or if they read this article (or a summary of it), will do they better?
If not, then we have an odd situation where a fairly simple specialized bot (eg Henny, or string-finder + Henny) can beat a SOTA general LLM! Surprising if true.
I'm going to test this when I get a chance. For now I'll register a prediction: for any given current frontier LLM and commonly-used sampler with a fixed context window using one of the approaches you described, the most common pick will still occur at least ⅔ of the time.
Never pass up a chance to share what might be my favourite moment from all of The SImpsons
https://youtu.be/b0SoKWLkmLU?si=CqWISTUakB--1lD7
I shared it in the post! :P