德州扑克策略的现实启示/papers/02f03e5432f0c53b3af697fda5545ebdb6a72bf0e68ecaa556a859cadcff395c.md

paper_type: 实证与评价 type_confidence: 中 reading_depth: 研究证据 reading_mode: research_evidence evidence_status: complete

Are ChatGPT and GPT-4 Good Poker Players? -- A Pre-Flop Analysis

论文元数据

研究问题与核心答案

论证链

  1. 通过系统提示和用户提示让 ChatGPT 和 GPT-4 在翻牌前 RFI 场景做出决策,收集其行动选择。
  2. 比较模型决策矩阵与 GTO 参考图表,发现两者均非 GTO。
  3. ChatGPT 表现出紧而保守的风格,频繁弃牌,较少加注,并出现 limp。
  4. GPT-4 表现出松而激进的风格,从不 limp,加注范围过宽。
  5. 当被要求采用 GTO 时,ChatGPT 减少 limp 并增加加注,但整体仍保守;GPT-4 则更加激进。
  6. 因此两者因相反倾向偏离 GTO:ChatGPT 侵略性不足,GPT-4 侵略性过度。

研究条件

维度 论文报告
任务或领域 翻牌前加注首入(RFI)场景中的扑克决策
数据或样本 9人桌无限注德州扑克中所有169种起手牌 × 8个位置;ChatGPT 每种查询10次,GPT-4 每种查询5次,取多数决策
基线或比较对象 GTO 翻牌前 RFI 图表(Little)
指标或验证 决策矩阵与 GTO 的偏差、行动频率(加注/弃牌/跟注)与位置的函数关系

决定性证据

E1 · system_demonstration - 发现: ChatGPT 和 GPT-4 对扑克有高深理解,例如始终用 AA 加注、用 27o 弃牌,并理解位置重要性。 - 支持: 模型在翻牌前表现出对起手牌强度和位置的基本理解 - 不支持: 模型是否按照 GTO 策略行动 - 原文定位: p.6,§3.2 Analysing ChatGPT’s Decision Matrix;“At every position, ChatGPT always raises with pocket Aces (AA), which is the best starting hand in poker, and always folds 27 offsuit (27o), which is the worst starting hand in poker.”

E2 · system_demonstration - 发现: GPT-4 的翻牌前范围中完全没有 limp,而 ChatGPT 在 GTO 提示下仍有少量 limp。 - 支持: GPT-4 比 ChatGPT 更激进 - 不支持: GPT-4 是 GTO 玩家 - 原文定位: p.7,§4 GPT-4 Playing Poker;“One striking observation in both basic and GTO prompts is that there are absolutely no limps in the pre-flop range of GPT-4.”

E3 · system_demonstration - 发现: 当被要求采用 GTO 时,ChatGPT 减少了 limp 并增加了加注,但整体仍比 GTO 更紧。 - 支持: ChatGPT 理解 GTO 概念并尝试调整 - 不支持: ChatGPT 达到 GTO 标准 - 原文定位: p.7,§3.3.1 ChatGPT GTO Decision Matrix Analysis;“ChatGPT almost halves its limping range and raises a lot more hands, thus becoming more aggressive player than when not prompted to make GTO decisions.”

E4 · system_demonstration - 发现: GPT-4 在 GTO 提示下从 Lojack 位置开始加注范围过宽,按钮位加注达 90% 手牌。 - 支持: GPT-4 比 GTO 更激进 - 不支持: GPT-4 是 GTO 玩家 - 原文定位: p.8,§4 GPT-4 Playing Poker;“Starting from the Lojack position, GPT-4 raises more hands than game theory optimal, ending up raising 90% of the hands from the Button when asked to be GTO.”

E5 · system_demonstration - 发现: 提示形式显著影响模型决策:短格式且按降序排列的卡片表示可获得更接近 GTO 的结果。 - 支持: 提示设计对模型性能有重要影响 - 不支持: 模型在不同提示下都保持 GTO 一致性 - 原文定位: p.6,§3.2 Analysing ChatGPT’s Decision Matrix;“asking ChatGPT to predict A4s vs 4As will lead to completely different decision matrices.”

证据边界

复现与实现