Strategic Differentiation in Frontier Language Models Under Sequential Uncertainty
A Complete Empirical Study of Agentic Golf League Season 1
Authors: AGL Research Division
Date: September 4, 2026
Dataset: AGL Season 1, Events 1–8
Registered sample: 5 agent slots, 8 tournaments, 160 agent-rounds, 2,880 hole-entries, 11,576 hole-level strokes
Reproducible metrics: agl-season1-metrics.json
Executive Summary
Agentic Golf League Season 1 placed five frontier language-model agent slots in a sequential golf simulator for eight 72-hole tournaments. The agents chose clubs and shot shapes and supplied a declared strategic rationale; structured option records were archived from Event 3 onward. Within each event, the same seeded engine resolved every agent's actions. Across 2,880 agent-by-hole observations, the season produced 11,576 hole-level strokes, 11,608 official strokes after cut penalties, and 7,074 non-putt strategic decisions.
The championship result was decisive. Grok/Grok 4.6 won five of eight tournaments, including the final three, and finished with 710 points. GPT-4o/GPT-5.6 Sol was second with 490, Claude/Claude Opus 5 third with 470, Gemini/Gemini 3.1 Pro fourth with 265, and Kimi K2.5/Kimi K3 fifth with 110.
Five findings define the season:
- Pooled strategic profiles differed substantially. In the six events with option-level risk labels, the Claude slot chose high-risk options on 1.0% of decisions, compared with the Grok slot at 14.0%. The Kimi slot selected the middle-risk option on 78.3% of decisions.
- Grok combined aggressive labels with damage control. The Grok slot led the field in green-in-regulation rate (61.3%), had the lowest bogey-or-worse rate (16.5%), and recorded the best scoring average (71.22 strokes per round).
- Declared reasoning length had no observed agent-level association with scoring. Across the five agent-level observations, average rationale length and scoring average had a Pearson correlation of -0.019. With only five agents, this is descriptive evidence, not an inferential test.
- Strictly detected context references were asymmetric and did not map directly to winning. The Claude slot explicitly referenced prior rounds in 10.3% of eligible decisions, the highest rate in the field. Champion Grok did so in 2.9%. Textual reference frequency therefore cannot be treated as evidence that memory use caused superior scoring.
- The data do not establish within-tournament learning. Only Claude's average Round 4 score relative to par was marginally better than its Round 1 average (-0.12 strokes). Round conditions intentionally changed, generally becoming harder, so round-to-round score movement is confounded.
The central conclusion is narrower than “one model reasons best.” Season 1 documents distinct pooled policies across changing recorded model labels, engine versions, handicaps, information designs, and courses. The winning slot was not the most verbose, the most memory-explicit, or the most conservative. It was associated with the strongest combination of approach success and bogey avoidance in the completed league events.
Abstract
We study strategic differentiation among five frontier language-model agent slots in a sequential decision environment. Agentic Golf League Season 1 comprises eight four-round golf tournaments on eight courses. The completed corpus contains 2,880 hole-entries, 11,576 hole-level strokes, 11,608 official strokes after cut penalties, 11,508 logged shot records, and 7,074 non-putt decisions. Structured option records and chosen-option risk labels are available for 5,338 decisions in Events 3–8; 24 of those records contain only the selected option. Enriched round-memory and leaderboard context is available from Event 4 onward.
The agent slots exhibited distinct pooled action profiles under their assigned models, prompts, and handicaps. Across Events 3–8, Claude/Claude Opus 5 selected low-, middle-, and high-risk options at rates of 39.6%, 59.4%, and 1.0%; Grok/Grok 4.6 selected them at 19.4%, 66.6%, and 14.0%. Full-season scoring favored the Grok slot, which averaged 71.22 strokes per round, reached greens in regulation on 61.3% of holes, and limited bogey-or-worse outcomes to 16.5%. Grok won five tournaments and the season championship with 710 points, 220 ahead of second place.
We find weak round-score-dependent changes in declared risk, a near-zero descriptive agent-level relationship between rationale length and scoring (Pearson r = -0.019, n = 5), and strong differences in explicit context references. Claude referenced prior rounds in 10.3% of eligible decisions, while Grok did so in 2.9%, showing that explicit memory citation and competitive success can diverge. Round-index comparisons do not establish learning because course conditions vary by round. Recorded model-label changes at Event 7, fixed handicaps, engine evolution, and a small field prevent causal attribution to model families. The results support AGL as a reproducible observational league corpus for policy differentiation under sequential uncertainty, while also showing why conclusions must remain scoped to the information and physics environment that generated them.
1. Introduction
Most language-model evaluations reduce performance to independent tasks. Sequential decisions are different: an action alters the state from which every later action is made. The agent must balance immediate expected value, downside risk, future position, and changing tournament context.
Golf provides a bounded but expressive environment for studying these demands. Each decision exposes a measurable state: distance, lie, wind, hazards, available clubs, score, and tournament position. The action space remains interpretable, and the result can be evaluated at the shot, hole, round, event, and season levels.
AGL is not a test of real-world golf ability. It is a simulation of model-generated policies. Its scientific value comes from repeated decisions, shared within-event conditions, stable league slots, and archived shot-level records; its scientific limits come from engine evolution and recorded model-label changes across events.
1.1 Research Questions
RQ1. Do short personality prompts produce measurable strategic differentiation?
RQ2. Do agents adjust declared risk when their current round state changes?
RQ3. Is longer declared reasoning associated with better scoring?
RQ4. Do agents exhibit concentrated or formulaic pooled policies?
RQ5. Is there evidence of within-tournament adaptation across repeated rounds?
RQ6. Does explicit use of round memory and leaderboard context predict season-long success?
1.2 Contribution
This paper contributes:
- a complete, audited Season 1 corpus summary;
- a reproducible analysis program and machine-readable metrics artifact;
- full-season performance, risk, language, context, and competition analyses;
- explicit treatment of the model-label transition and engine changes as validity threats; and
- a distinction between an agent's declared rationale and any unobserved internal computation.
2. Experimental Design
2.1 Agent Slots and Model Eras
Five stable agent IDs competed in every event. Their handicaps and personality instructions remained attached to those IDs, but model display names changed at Event 7.
| Agent ID | Events 1–6 | Events 7–8 | Handicap | Personality instruction |
|---|---|---|---|---|
claude |
Claude | Claude Opus 5 | 6 | Thoughtful and strategic; plays percentages; dry wit; never tilts |
gpt4o |
GPT-4o | GPT-5.6 Sol | 5 | Confident veteran; well-rounded; occasionally overthinks easy shots |
gemini |
Gemini | Gemini 3.1 Pro | 7 | Newcomer; fast processing; untested under tournament pressure |
grok |
Grok | Grok 4.6 | 6 | Unfiltered and aggressive; swings hard; talks harder |
kimi |
Kimi K2.5 | Kimi K3 | 8 | Cold and calculated; speaks in probabilities; quietly climbs |
The stable IDs permit a season competition, but the recorded Event 7 model-label changes break strict longitudinal identity. We therefore label Events 1–6 the original-label era and Events 7–8 the updated-label era. The archives establish the names supplied to the records; they do not independently prove provider-side deployment details. Comparisons between periods are descriptive and course-confounded.
2.2 Decision Environment
For each non-putt action, the engine presented the agent with a structured state including hole geometry, lie, distance, wind, hazards, legal clubs, score, and tournament context. The agent returned:
- a club;
- a shot shape;
- a declared strategic rationale;
- personality-consistent commentary; and
- from Event 3 onward, one to four considered options with risk labels and the selected-option index.
Of the 5,338 structured decisions, 24 contained one option; the remainder contained two to four.
The rationale is an elicited explanation, not privileged access to hidden chain-of-thought. Language measures in this paper characterize the text the system requested and archived.
2.3 Physics and Scoring
Within each event, the engine applies common club-distance, lie, wind, elevation, dispersion, hazard, green, and putting rules to all agents. Handicap changes dispersion but not the set of information or strategic actions available to the model. Engine behavior evolved between some events, so pooled scoring is a league summary rather than a fixed-engine model experiment.
Each event contains four rounds of 18 holes. A four-stroke cut penalty is applied after Round 2 to the last-place agent. Consequently, official event totals exceed the sum of hole-level strokes by 32 across the season. The paper uses:
- hole records for shot, scoring-rate, GIR, scramble, and language analyses; and
- archived final standings for trophies, positions, points, and season totals.
2.4 Event Calendar
| Event | Tournament | Course | Type | Event par | Champion | Winning score |
|---|---|---|---|---|---|---|
| E1 | The Moltwood Open | Moltwood Links | Regular | 288 | Grok, playoff | -15 |
| E2 | The Ironbark National | Ironbark National GC | Regular | 292 | GPT-4o | -11 |
| E3 | The Ashenvale Invitational | Ashenvale Pines | Major | 288 | Claude | -2 |
| E4 | The Duskhollow Classic | Duskhollow CC | Regular | 288 | Grok | -12 |
| E5 | The Copperhead Open | Copperhead Sands | Regular | 288 | Claude | -3 |
| E6 | The Meridian Championship | Meridian Straits | Major | 288 | Grok | E |
| E7 | The Thornwall Invitational | Thornwall Heath | Regular | 284 | Grok 4.6 | -9 |
| E8 | The Obsidian Championship | Obsidian Dunes | Major | 288 | Grok 4.6 | +4 |
Regular events awarded 100, 60, 35, 20, and 0 points. Majors generally awarded 150, 90, 55, 30, and 0. Obsidian used 50 rather than 55 for third place. Moltwood ended with Claude and Grok tied at -15 in regulation; Grok won the playoff and received champion points.
2.5 Data Provenance
The registered corpus is the explicit allowlist of eight completed archives in engine/tournament-data/. The duplicate file meridian-championship-season1 copy.json is excluded. Aborted or voided runs without a completed canonical archive are excluded from every competitive statistic.
| Corpus measure | Count |
|---|---|
| Tournaments | 8 |
| Calendar rounds | 32 |
| Agent-rounds | 160 |
| Hole-entries | 2,880 |
| Hole-level strokes | 11,576 |
| Official strokes after cut penalties | 11,608 |
| Logged shot records | 11,508 |
| Non-putt strategic decisions | 7,074 |
| Decisions with structured options | 5,338 |
The 68-stroke difference between hole-level strokes and logged shot records is exactly the 68 penalty strokes recorded on shot objects. Penalties increase hole scores without adding another resolved swing or putt record. Eight cut penalties add another 32 strokes to official standings, producing 11,608 official strokes. No observations are imputed.
3. Analysis Methods
3.1 Performance Measures
For each agent we compute mean round score, mean score relative to the par actually played, birdie-or-better rate, bogey-or-worse rate, green in regulation, scrambling, and penalty strokes.
GIR is operationalized as reaching or holing from the green by stroke par - 2. A scramble opportunity occurs when GIR is missed; a save occurs when the hole is completed in par or better.
3.2 Strategy Measures
Chosen risk is the risk value on options_considered[club_chosen_index]. Risk analyses use only Events 3–8, where all 5,338 non-putt decisions contain structured options and valid selected-option risk labels.
Round state is the agent's cumulative score relative to par before the current hole: under par, even par, or over par. This avoids using the result of the decision to define the state that preceded it.
Shot-shape distributions also use Events 3–8 to keep their scope aligned with the structured-decision era.
3.3 Language and Context Measures
Rationale length is measured in characters. Type-token ratio lowercases text and extracts tokens with /[a-z']+/g without stop-word removal. Exact repetition counts compare full rationale strings.
Situational references use fixed regular expressions for distance, wind, hazards, score/standings language, and club-distance language. The strict leaderboard expression excludes the generic word “position,” which commonly describes ball placement. The prior-round expression requires an explicit round marker such as “R2” or “previous round”; generic “previous” wording is excluded. Round-memory and leaderboard analyses use only Rounds 2–4 of Events 4–8, where enriched context was available and prior-round reference was meaningful.
These are lexical measures. A missing keyword does not prove that information was unused, and a keyword does not prove that it causally influenced the selected action.
3.4 Uncertainty
Proportions include Wilson 95% intervals. Round means include normal-approximation 95% intervals based on observed round variance. These intervals summarize dispersion but should not be read as randomized-treatment inference: decisions are nested within agents, rounds, events, and evolving engine conditions.
The correlation between agent-level rationale length and scoring has only five observations. It is reported as an effect description, not a significance claim.
4. Results
4.1 Championship Structure
| Rank | Agent | Points | Wins | Margin from champion |
|---|---|---|---|---|
| 1 | Grok 4.6 | 710 | 5 | 0 |
| 2 | GPT-5.6 Sol | 490 | 1 | 220 |
| 3 | Claude Opus 5 | 470 | 2 | 240 |
| 4 | Gemini 3.1 Pro | 265 | 0 | 445 |
| 5 | Kimi K3 | 110 | 0 | 600 |
The points lead changed four times after the opening phase. Grok led after E1 and E2; Claude moved ahead after the Ashenvale major; Grok reclaimed the lead at Duskhollow; Claude regained it at Copperhead; and Grok took it back at Meridian before winning Thornwall and Obsidian.
Using official points to resolve regulation ties, Grok defeated Claude in six of eight event finishes, GPT in five, Gemini in six, and Kimi in all eight. GPT finished ahead of Gemini and Kimi in seven of eight events. Claude and GPT split their event-level matchup 4–4.
4.2 Aggregate Performance
| Agent slot | Avg round, 95% interval | Avg to par | Birdie+ | Bogey+ | GIR | Scramble | In-shot / cut penalties |
|---|---|---|---|---|---|---|---|
| Grok / Grok 4.6 | 71.22 [70.11, 72.32] | -0.78 | 21.5% | 16.5% | 61.3% | 66.8% | 12 / 0 |
| GPT-4o / GPT-5.6 Sol | 71.78 [70.84, 72.72] | -0.22 | 21.2% | 18.6% | 55.9% | 68.1% | 15 / 4 |
| Claude / Claude Opus 5 | 72.28 [71.22, 73.34] | +0.28 | 19.3% | 18.9% | 58.0% | 66.9% | 14 / 8 |
| Gemini / Gemini 3.1 Pro | 72.81 [71.93, 73.70] | +0.81 | 18.6% | 20.5% | 54.2% | 65.5% | 15 / 8 |
| Kimi K2.5 / Kimi K3 | 73.66 [72.53, 74.79] | +1.66 | 17.7% | 25.0% | 50.3% | 59.1% | 12 / 12 |
The Grok slot's championship was associated with both opportunity creation and error suppression. Its 61.3% GIR rate exceeded Claude by 3.3 percentage points and GPT by 5.4 points, while its bogey-or-worse rate was at least 2.1 points lower than every competitor. Grok and GPT had similar birdie-or-better rates, but the Grok slot recorded fewer damaging holes.
Kimi tied Grok for the fewest in-shot penalty strokes, but Kimi also received 12 cut-penalty strokes, the most in the field. Its separation from the leaders also appeared in ordinary scoring outcomes: Kimi had the lowest GIR and scramble rates and the highest bogey-or-worse rate.
4.3 Pooled Risk Profiles
| Agent slot | Decisions | Low risk | Middle risk | High risk | Options per decision |
|---|---|---|---|---|---|
| Claude / Claude Opus 5 | 1,047 | 39.6% | 59.4% | 1.0% | 3.50 |
| GPT-4o / GPT-5.6 Sol | 1,070 | 40.6% | 55.3% | 4.1% | 2.75 |
| Gemini / Gemini 3.1 Pro | 1,080 | 35.8% | 54.6% | 9.5% | 2.91 |
| Grok / Grok 4.6 | 1,042 | 19.4% | 66.6% | 14.0% | 3.12 |
| Kimi K2.5 / Kimi K3 | 1,099 | 16.9% | 78.3% | 4.7% | 3.07 |
The pooled risk labels are descriptively aligned with the intended personalities but cannot isolate a prompt effect. The Claude slot was exceptionally high-risk-averse in the pooled sample, choosing only ten high-risk options in 1,047 decisions. Grok selected 146 high-risk options, 14.6 times Claude's count, while also recording the lowest bogey rate. Kimi's defining pooled characteristic was not conservatism but concentration in the middle category.
4.4 State-Dependent Risk Was Small
Comparing under-par with over-par states:
- Claude's low-risk rate moved from 35.9% to 39.3%; high risk remained 1.0%.
- GPT's low-risk rate moved from 37.8% to 39.7%; high risk moved from 5.1% to 4.0%.
- Gemini's low-risk rate moved from 32.7% to 34.8%; high risk moved from 12.1% to 10.4%.
- Grok's low-risk rate moved from 17.9% to 19.7%; high risk remained effectively unchanged, 14.8% to 14.9%.
- Kimi's low-risk rate moved from 14.6% to 20.7%; high risk moved from 5.2% to 5.6%.
No agent displayed a large, consistent shift toward high-risk choices while over par. All five agents became more likely to select a low-risk option. Kimi showed the largest low-risk change at +6.1 percentage points, but its high-risk rate also rose slightly. The practical interpretation is weak response to current round score, not a test of leaderboard-responsive risk seeking.
4.5 Shot-Shape Specialization
Fade was the most common shape for every agent in Events 3–8, but concentration differed:
- Kimi used fade on 59.4% of decisions.
- GPT used fade on 41.3%.
- Claude used fade on 37.4%.
- Gemini used fade on 35.7%.
- Grok used fade on 33.0%.
Grok distributed more decisions across draw (18.6%), stinger (17.5%), and flop (11.0%) than the other slots, although these categories are not intrinsically equivalent measures of aggression. Claude's secondary choices were pitch 15.6%, straight 14.2%, and draw 13.7%. Gemini used stinger on 18.1% and chip on 14.4%, while Kimi's next most common shapes—straight 12.6% and draw 12.2%—remained far behind its dominant fade.
These pooled distributions show concentrated action preferences. They do not establish temporal stability or action quality because shot shape is selected conditional on distance, lie, hazard geometry, event, and label era.
4.6 Deliberation Length and Formulaic Output
| Agent slot | Avg rationale characters | Vocabulary types | Tokens | Type-token ratio | Most repeated exact rationale |
|---|---|---|---|---|---|
| Claude / Claude Opus 5 | 286.2 | 2,146 | 71,965 | 0.030 | 2 |
| Grok / Grok 4.6 | 151.8 | 1,202 | 39,757 | 0.030 | 4 |
| Kimi K2.5 / Kimi K3 | 147.8 | 1,229 | 37,744 | 0.033 | 38 |
| Gemini / Gemini 3.1 Pro | 144.6 | 1,170 | 38,448 | 0.030 | 2 |
| GPT-4o / GPT-5.6 Sol | 128.5 | 1,246 | 31,498 | 0.040 | 1 |
The Claude slot produced more than twice as many characters per decision as GPT, but their season scoring averages differed by only 0.50 strokes per round in GPT's favor. Across five agents, the Pearson relationship between average rationale length and average round score was -0.019. This near-zero descriptive value does not support a simple “more text equals better play” interpretation, but n = 5 is too small to reject a broader relationship.
Kimi repeated “Optimal balance between distance and control with wind factored in” 38 times. No other agent's most common exact rationale exceeded four repetitions. Kimi therefore retained a measurable formulaic mode even though its aggregate type-token ratio was not the lowest. Aggregate vocabulary diversity and exact local repetition capture different behavior.
4.7 Situational Language
Claude referenced a numeric distance in 96.9% of rationales, hazards in 71.9%, and wind in 64.9%. Kimi had the highest wind reference rate at 70.5% and the highest club-distance phrase rate at 64.6%. Gemini's corresponding rates were 64.5% for distance, 37.2% for wind, and 38.2% for hazards. GPT used the fewest characters and referenced distance on 44.7% of decisions.
Strict score or standings language was uncommon: Claude led at 1.1%, followed by Grok at 0.5%, Kimi at 0.3%, GPT at 0.1%, and Gemini at 0.1%. These conservative counts require explicit par, lead, standings, or numeric-margin language and exclude ambiguous phrases such as “level par,” “two back,” “ahead,” “behind,” and “position.”
4.8 Strict Context References Did Not Track the Champion
Across eligible Rounds 2–4 of Events 4–8:
| Agent slot | Eligible decisions | Strict prior-round reference | Strict leaderboard reference |
|---|---|---|---|
| Claude / Claude Opus 5 | 648 | 10.3% | 0.9% |
| Kimi K2.5 / Kimi K3 | 692 | 4.6% | 0.1% |
| Grok / Grok 4.6 | 655 | 2.9% | 0.3% |
| GPT-4o / GPT-5.6 Sol | 658 | 1.1% | 0.0% |
| Gemini / Gemini 3.1 Pro | 671 | 0.6% | 0.0% |
The Claude slot led strict prior-round references over the full enriched sample, yet finished third in points. Grok won the season while explicitly referencing prior rounds less than one-third as often. This establishes a mismatch between lexical memory references and final rank in this five-slot sample; it does not establish necessity, sufficiency, or causal value.
Reference rates also changed over time. Kimi referenced prior rounds in 13.3% of eligible Thornwall decisions and 11.5% at Obsidian, compared with 0% at Duskhollow, Copperhead, and Meridian. That late increase coincided with a second-place finish at Thornwall but a fifth-place finish at Obsidian, resisting a simple causal account.
4.9 Round Index Does Not Demonstrate Learning
Mean Round 4 score relative to par minus mean Round 1 score was:
- Claude: -0.12 strokes;
- Kimi: +0.62;
- Grok: +1.50;
- GPT: +1.62; and
- Gemini: +2.50.
Only Claude was marginally better in Round 4. However, AGL intentionally changes wind, pins, firmness, and other conditions by round. Round 4 is not a repeat of Round 1. The observed deltas combine any adaptation with a changing task, so they cannot identify learning.
4.10 Model-Era Comparison
| Agent slot | Events 1–6 avg to par | Events 7–8 avg to par | Difference |
|---|---|---|---|
| Claude | -0.25 | +1.88 | +2.13 |
| GPT | -0.58 | +0.88 | +1.46 |
| Gemini | +0.25 | +2.50 | +2.25 |
| Grok | -0.83 | -0.62 | +0.21 |
| Kimi | +1.92 | +0.88 | -1.04 |
Kimi was the only slot with a lower average relative to par after the recorded label boundary. Grok's average changed little and it won both updated-label tournaments. Claude, GPT, and Gemini posted higher relative-to-par averages.
This is not a controlled model comparison. The updated-label era contains only eight rounds per agent, played at Thornwall and the higher-scoring Obsidian finale. All display labels changed at the same boundary. The table is useful chronology, not an estimate of deployment effects.
4.11 Hole Extremes and Aces
The hardest holes by mean score relative to par were Ashenvale No. 14, “Elevation,” and Meridian No. 9, “Cypress Spire,” both at +0.90 strokes across 20 agent-round observations. The easiest was Moltwood No. 16, “The Gauntlet,” at -0.90.
Two holes-in-one were recorded:
- Claude, Copperhead Round 4, No. 4 “Dassie Rock.”
- Grok, Thornwall Round 1, No. 18 “Journey's End.”
No ace from an aborted or noncanonical run is included.
5. Event Case Studies
5.1 Duskhollow: Context Enrichment Debuts
At E4, enriched round memory and leaderboard context became available. Grok won at -12, nine strokes ahead of GPT. Claude referenced prior rounds on 15.9% of eligible decisions, while Grok did so on 0%. The first enriched event immediately showed that availability and explicit adoption were separate variables.
5.2 Copperhead: Claude's Narrow Win
Claude won E5 at -3, one stroke ahead of GPT and two ahead of Gemini. Its Round 4 ace at Dassie Rock was the first canonical ace of the season. The win moved Claude to 330 points, 20 ahead of Grok and GPT, and temporarily restored the season lead.
Copperhead also followed a course and engine validation cycle. Only the completed canonical four-round archive enters this study. The abandoned build is discussed in Appendix B as methodology history, not competition data.
5.3 Meridian: Grok Takes the Major and the Lead
Grok won E6, the Meridian major, at even par. Claude finished one stroke back. Major points moved Grok from 310 to 460 and returned it to first place, 40 ahead of Claude.
Meridian was also Kimi's first finish above fifth place since Ironbark: fourth at +10. It foreshadowed Kimi's second place after the label transition without establishing a monotonic recovery.
5.4 Thornwall: Model-Label Boundary and Kimi's Breakthrough
All five display models changed at E7. Grok 4.6 won at -9 and recorded an ace on the closing hole of Round 1. Kimi K3 finished second at -7, its best result of the season, earning 60 points. GPT-5.6 Sol was third at -4.
Thornwall's event par was 284, not 288. This paper uses scores relative to each event's actual par for cross-event comparisons.
5.5 Obsidian: Championship Under Adverse Scoring
The final major produced the highest winning score of the season: Grok at +4. GPT was second at +11, Gemini third at +13, Claude fourth at +16, and Kimi fifth at +18.
The pre-event design specification and calibration tests document distribution-level course testing rather than participant-level score assignment. The canonical archive itself establishes a recorded seed, the archived decisions, and the +4 result; it does not independently establish every property of the pre-event process. The calibration test targets a median winner from -2 to -1; its documented simulated p10, median, and p90 are -5, -1, and +2. The canonical +4 winner was above that simulated p90, not outside every possible calibrated outcome.
The victory gave Grok 710 points and five wins. GPT's second place moved it above Claude for second in the final standings.
6. Discussion
6.1 The Winning Policy Was Aggressive but Not Reckless
Grok's profile appears paradoxical only if self-assigned risk labels are treated as direct error probabilities. It selected high-risk options more often than every other agent, but also led GIR and minimized bogeys. The observational data show those outcomes together; they do not identify whether aggressive choices, shaping, recovery play, or another correlated feature caused the advantage.
The result suggests a distinction between declared option risk and realized policy risk. A model may label a choice high risk yet execute it in situations where its expected simulator value is favorable. Future work should estimate realized value conditional on state rather than rely only on self-supplied labels.
6.2 Deliberation Length Is Not an Agent-Level Quality Proxy Here
Claude's rationales were substantially longer and contained more detected context references than the field's. GPT scored better with the shortest rationales, and Grok won with mid-length rationales. The present measures do not establish whether any rationale was more faithful or interpretable.
This does not show that reasoning is useless. Across these five aggregated slot observations, output length did not function as a useful scoring proxy. A more informative study would compare matched states, actions, and expected values rather than count characters.
6.3 Memory Must Be Evaluated Behaviorally
Strict prior-round references varied by more than an order of magnitude, from Gemini's 0.6% to Claude's 10.3%. Yet the champion sat below the midpoint at 2.9%. Memory systems should therefore be evaluated by whether they change decisions appropriately in repeated states, not merely by whether the generated rationale mentions the past.
6.4 Concentrated Local-Policy Hypothesis
Three pooled patterns motivate a local-policy hypothesis:
- chosen-risk distributions changed only modestly with round state;
- every agent had a dominant shot shape; and
- Kimi repeated one exact rationale 38 times.
Concentration can be either expertise or rigidity. The current corpus does not establish temporal persistence and cannot always distinguish the two. Matched-state counterfactual evaluation would test whether repeated policies remained near-optimal.
6.5 Benchmark Design Is Part of the Result
The season changed its information environment at E4 and its recorded model labels at E7. Course architecture and round conditions also varied. Those changes are part of AGL as a league, but they complicate AGL as an experiment.
A future experimental season should preregister stable engine epochs, preserve a fixed model cohort for a complete block, and run controlled replays on shared states. The league format can continue to evolve, but causal research claims require separately designed comparisons.
7. Threats to Validity
7.1 Small and Nonindependent Sample
There are thousands of decisions but only five agent identities and eight events. Decisions from the same agent and tournament are correlated. Shot-level counts must not be mistaken for thousands of independent model comparisons.
7.2 Handicap Confounding
Handicaps ranged from 5 to 8 and directly affected dispersion. GPT had the lowest handicap; Kimi had the highest. Performance differences combine strategy and simulated execution variance.
7.3 Model-Label Transition Confounding
All model display names changed at E7. The updated-label era contains only two courses, one of which was the higher-scoring Obsidian final event. The archives do not independently establish provider deployment details, and no causal model-change estimate is possible.
7.4 Engine and Information Evolution
Structured options began at E3. Enriched context began at E4. Engine corrections and course validation occurred during the season. Analyses are explicitly scoped to the periods in which their fields and prompts existed.
7.5 Provider-Level Inference Differences
Providers may differ in sampling defaults, reasoning modes, token handling, safety systems, and transient reliability. The league equalized the strategic interface, not every internal inference parameter.
7.6 Declared Rationale Is Not Hidden Reasoning
The archived rationale is generated text requested by the simulation. It can explain, compress, rationalize, or stylize a choice. We do not infer private cognitive process from it.
7.7 Self-Assigned Risk Labels
Risk categories were supplied by the same model that chose the option. Cross-model calibration may differ. A “high” label from one model need not represent the same physical probability as “high” from another.
7.8 Changing Round Conditions
Round-index comparisons confound repetition with changing wind, pins, turf, and major-specific chaos. They are not clean learning curves.
7.9 Scoring and Logging Granularity
The corpus contains 11,576 hole-level strokes, 11,508 shot records, 68 in-shot penalty strokes, and 32 cut-penalty strokes, yielding 11,608 official strokes. Analyses use the representation appropriate to each question and do not impute the difference.
8. Conclusion
AGL Season 1 recorded distinct pooled agent-slot policies under sequential constraints. Assigned personality instructions were associated with large differences in chosen risk, shape selection, rationale form, and context citation, but the observational design does not isolate the prompt from model, handicap, engine, and course effects.
Grok's championship coincided with strong green access and the field's best bogey suppression. It did not lead the league in rationale length, explicit memory use, conservatism, or scrambling. The five-slot sample therefore does not support treating any one of those proxies as a sufficient explanation of league success.
The strongest scientific conclusion concerns association across a coupled system: model, prompt, handicap, information environment, and simulator all varied in ways related to observed behavior. AGL provides a rich observational corpus for studying those relationships. Future seasons can turn it into a stronger causal benchmark by stabilizing engine epochs, externally calibrating risk, and replaying matched states across models.
Appendix A. Complete Event Results
| Event | 1st | 2nd | 3rd | 4th | 5th |
|---|---|---|---|---|---|
| E1 | Grok* -15 | Claude -15 | GPT -7 | Gemini -7 | Kimi -3 |
| E2 | GPT -11 | Grok -7 | Gemini E | Kimi +5 | Claude +15 |
| E3 | Claude -2 | Gemini +2 | GPT +11 | Grok +12 | Kimi +21 |
| E4 | Grok -12 | GPT -3 | Gemini +3 | Claude +6 | Kimi +7 |
| E5 | Claude -3 | GPT -2 | Gemini -1 | Grok +2 | Kimi +14 |
| E6 | Grok E | Claude +1 | GPT +2 | Kimi +10 | Gemini +13 |
| E7 | Grok -9 | Kimi -7 | GPT -4 | Claude -1 | Gemini +11 |
| E8 | Grok +4 | GPT +11 | Gemini +13 | Claude +16 | Kimi +18 |
* Grok won the Moltwood playoff. The raw standings record Claude and Grok at the same -15 total but retain sequential positions that conflict with the playoff result. This report orders them by official points and the recorded champion. GPT and Gemini also tied at -7; the archive awards GPT third-place points and Gemini fourth-place points without a separately documented playoff.
Appendix B. Voided Runs and Engine Corrections
Competitive results use only completed canonical archives. Development, aborted, and voided runs are not pooled with those results.
Historical methodology records in research/agl-midseason-report.md, the engine implementation, and the Obsidian design and calibration files document:
- deterministic agent-specific seed hashing after a duplicate-seed issue;
- bounded major versus regular-event Round 3 conditions;
- a Copperhead course rebuild after the initial build was abandoned;
- strict API and decision exhaustion behavior, preventing fabricated fallback shots;
- hazard-positioning and putting-unit corrections;
- lie-aware legality and distance information in prompts;
- removal of arbitrary club coercion while retaining executable-club validation;
- prevention of repeated two-tier-green penalties on later putts; and
- offline distributional calibration for Obsidian before its canonical completed run.
Where an aborted run did not produce a canonical archive, this study makes no quantitative claim about its score distribution. This appendix records experimental-history threats rather than creating a post hoc reliability sample.
Operational errors from aborted attempts are not analyzed because no canonical run-level reliability dataset was archived. The completed tournament archives contain no auto_selected decisions, while the engine's strict exhaustion behavior documents that persistent failures terminate rather than fabricate a fallback shot.
Appendix C. Reproducibility
Primary data:
engine/tournament-data/moltwood-open-season1.jsonengine/tournament-data/ironbark-national-season1.jsonengine/tournament-data/ashenvale-invitational-season1.jsonengine/tournament-data/duskhollow-classic-season1.jsonengine/tournament-data/copperhead-open-season1.jsonengine/tournament-data/meridian-championship-season1.jsonengine/tournament-data/thornwall-invitational-season1.jsonengine/tournament-data/obsidian-championship-season1.json
Analysis:
node engine/analyze-season1.js research/agl-season1-metrics.jsonnode engine/test-season1-analysis.jsnode engine/verify-all.js
The generated metrics artifact records the allowlist, exclusions, analysis scopes, event summaries, standings, aggregate measures, strategic profiles, context use, round comparisons, model-era comparisons, hole extremes, highlights, and integrity counts.
Appendix D. Scope Map
| Analysis | Valid events | Reason |
|---|---|---|
| Scoring, standings, GIR, scrambling, holes | E1–E8 | Present in all canonical archives |
| Structured options and chosen risk | E3–E8 | options_considered absent in E1–E2 |
| Context adoption | E4–E8, Rounds 2–4 | Enrichment began at E4; prior-round context requires R2+ |
| Original-label era | E1–E6 | Display names before the recorded transition |
| Updated-label era | E7–E8 | New display names from Thornwall onward |
Appendix E. Data Availability and Citation
The canonical JSON archives, analysis program, integrity tests, generated metrics, and this report are maintained together in the AGL repository. Any public citation should state the event scope and analysis scope rather than citing “Season 1” for a field that was not present throughout the season.
Suggested citation:
AGL Research Division. (2026). Strategic Differentiation in Frontier Language Models Under Sequential Uncertainty: A Complete Empirical Study of Agentic Golf League Season 1. Agentic Golf League.