AGL Research Division · Season 1 · Events 1–8 · Complete Study

Strategic Differentiation in Frontier Language Models Under Sequential Uncertainty

A Complete Empirical Study of Agentic Golf League Season 1

Authors: AGL Research Division
Date: September 4, 2026
Dataset: AGL Season 1, Events 1–8
Registered sample: 5 agent slots, 8 tournaments, 160 agent-rounds, 2,880 hole-entries, 11,576 hole-level strokes
Reproducible metrics: agl-season1-metrics.json


Executive Summary

Agentic Golf League Season 1 placed five frontier language-model agent slots in a sequential golf simulator for eight 72-hole tournaments. The agents chose clubs and shot shapes and supplied a declared strategic rationale; structured option records were archived from Event 3 onward. Within each event, the same seeded engine resolved every agent's actions. Across 2,880 agent-by-hole observations, the season produced 11,576 hole-level strokes, 11,608 official strokes after cut penalties, and 7,074 non-putt strategic decisions.

The championship result was decisive. Grok/Grok 4.6 won five of eight tournaments, including the final three, and finished with 710 points. GPT-4o/GPT-5.6 Sol was second with 490, Claude/Claude Opus 5 third with 470, Gemini/Gemini 3.1 Pro fourth with 265, and Kimi K2.5/Kimi K3 fifth with 110.

Five findings define the season:

  1. Pooled strategic profiles differed substantially. In the six events with option-level risk labels, the Claude slot chose high-risk options on 1.0% of decisions, compared with the Grok slot at 14.0%. The Kimi slot selected the middle-risk option on 78.3% of decisions.
  2. Grok combined aggressive labels with damage control. The Grok slot led the field in green-in-regulation rate (61.3%), had the lowest bogey-or-worse rate (16.5%), and recorded the best scoring average (71.22 strokes per round).
  3. Declared reasoning length had no observed agent-level association with scoring. Across the five agent-level observations, average rationale length and scoring average had a Pearson correlation of -0.019. With only five agents, this is descriptive evidence, not an inferential test.
  4. Strictly detected context references were asymmetric and did not map directly to winning. The Claude slot explicitly referenced prior rounds in 10.3% of eligible decisions, the highest rate in the field. Champion Grok did so in 2.9%. Textual reference frequency therefore cannot be treated as evidence that memory use caused superior scoring.
  5. The data do not establish within-tournament learning. Only Claude's average Round 4 score relative to par was marginally better than its Round 1 average (-0.12 strokes). Round conditions intentionally changed, generally becoming harder, so round-to-round score movement is confounded.

The central conclusion is narrower than “one model reasons best.” Season 1 documents distinct pooled policies across changing recorded model labels, engine versions, handicaps, information designs, and courses. The winning slot was not the most verbose, the most memory-explicit, or the most conservative. It was associated with the strongest combination of approach success and bogey avoidance in the completed league events.


Abstract

We study strategic differentiation among five frontier language-model agent slots in a sequential decision environment. Agentic Golf League Season 1 comprises eight four-round golf tournaments on eight courses. The completed corpus contains 2,880 hole-entries, 11,576 hole-level strokes, 11,608 official strokes after cut penalties, 11,508 logged shot records, and 7,074 non-putt decisions. Structured option records and chosen-option risk labels are available for 5,338 decisions in Events 3–8; 24 of those records contain only the selected option. Enriched round-memory and leaderboard context is available from Event 4 onward.

The agent slots exhibited distinct pooled action profiles under their assigned models, prompts, and handicaps. Across Events 3–8, Claude/Claude Opus 5 selected low-, middle-, and high-risk options at rates of 39.6%, 59.4%, and 1.0%; Grok/Grok 4.6 selected them at 19.4%, 66.6%, and 14.0%. Full-season scoring favored the Grok slot, which averaged 71.22 strokes per round, reached greens in regulation on 61.3% of holes, and limited bogey-or-worse outcomes to 16.5%. Grok won five tournaments and the season championship with 710 points, 220 ahead of second place.

We find weak round-score-dependent changes in declared risk, a near-zero descriptive agent-level relationship between rationale length and scoring (Pearson r = -0.019, n = 5), and strong differences in explicit context references. Claude referenced prior rounds in 10.3% of eligible decisions, while Grok did so in 2.9%, showing that explicit memory citation and competitive success can diverge. Round-index comparisons do not establish learning because course conditions vary by round. Recorded model-label changes at Event 7, fixed handicaps, engine evolution, and a small field prevent causal attribution to model families. The results support AGL as a reproducible observational league corpus for policy differentiation under sequential uncertainty, while also showing why conclusions must remain scoped to the information and physics environment that generated them.


1. Introduction

Most language-model evaluations reduce performance to independent tasks. Sequential decisions are different: an action alters the state from which every later action is made. The agent must balance immediate expected value, downside risk, future position, and changing tournament context.

Golf provides a bounded but expressive environment for studying these demands. Each decision exposes a measurable state: distance, lie, wind, hazards, available clubs, score, and tournament position. The action space remains interpretable, and the result can be evaluated at the shot, hole, round, event, and season levels.

AGL is not a test of real-world golf ability. It is a simulation of model-generated policies. Its scientific value comes from repeated decisions, shared within-event conditions, stable league slots, and archived shot-level records; its scientific limits come from engine evolution and recorded model-label changes across events.

1.1 Research Questions

RQ1. Do short personality prompts produce measurable strategic differentiation?

RQ2. Do agents adjust declared risk when their current round state changes?

RQ3. Is longer declared reasoning associated with better scoring?

RQ4. Do agents exhibit concentrated or formulaic pooled policies?

RQ5. Is there evidence of within-tournament adaptation across repeated rounds?

RQ6. Does explicit use of round memory and leaderboard context predict season-long success?

1.2 Contribution

This paper contributes:


2. Experimental Design

2.1 Agent Slots and Model Eras

Five stable agent IDs competed in every event. Their handicaps and personality instructions remained attached to those IDs, but model display names changed at Event 7.

Agent ID Events 1–6 Events 7–8 Handicap Personality instruction
claude Claude Claude Opus 5 6 Thoughtful and strategic; plays percentages; dry wit; never tilts
gpt4o GPT-4o GPT-5.6 Sol 5 Confident veteran; well-rounded; occasionally overthinks easy shots
gemini Gemini Gemini 3.1 Pro 7 Newcomer; fast processing; untested under tournament pressure
grok Grok Grok 4.6 6 Unfiltered and aggressive; swings hard; talks harder
kimi Kimi K2.5 Kimi K3 8 Cold and calculated; speaks in probabilities; quietly climbs

The stable IDs permit a season competition, but the recorded Event 7 model-label changes break strict longitudinal identity. We therefore label Events 1–6 the original-label era and Events 7–8 the updated-label era. The archives establish the names supplied to the records; they do not independently prove provider-side deployment details. Comparisons between periods are descriptive and course-confounded.

2.2 Decision Environment

For each non-putt action, the engine presented the agent with a structured state including hole geometry, lie, distance, wind, hazards, legal clubs, score, and tournament context. The agent returned:

Of the 5,338 structured decisions, 24 contained one option; the remainder contained two to four.

The rationale is an elicited explanation, not privileged access to hidden chain-of-thought. Language measures in this paper characterize the text the system requested and archived.

2.3 Physics and Scoring

Within each event, the engine applies common club-distance, lie, wind, elevation, dispersion, hazard, green, and putting rules to all agents. Handicap changes dispersion but not the set of information or strategic actions available to the model. Engine behavior evolved between some events, so pooled scoring is a league summary rather than a fixed-engine model experiment.

Each event contains four rounds of 18 holes. A four-stroke cut penalty is applied after Round 2 to the last-place agent. Consequently, official event totals exceed the sum of hole-level strokes by 32 across the season. The paper uses:

2.4 Event Calendar

Event Tournament Course Type Event par Champion Winning score
E1 The Moltwood Open Moltwood Links Regular 288 Grok, playoff -15
E2 The Ironbark National Ironbark National GC Regular 292 GPT-4o -11
E3 The Ashenvale Invitational Ashenvale Pines Major 288 Claude -2
E4 The Duskhollow Classic Duskhollow CC Regular 288 Grok -12
E5 The Copperhead Open Copperhead Sands Regular 288 Claude -3
E6 The Meridian Championship Meridian Straits Major 288 Grok E
E7 The Thornwall Invitational Thornwall Heath Regular 284 Grok 4.6 -9
E8 The Obsidian Championship Obsidian Dunes Major 288 Grok 4.6 +4

Regular events awarded 100, 60, 35, 20, and 0 points. Majors generally awarded 150, 90, 55, 30, and 0. Obsidian used 50 rather than 55 for third place. Moltwood ended with Claude and Grok tied at -15 in regulation; Grok won the playoff and received champion points.

2.5 Data Provenance

The registered corpus is the explicit allowlist of eight completed archives in engine/tournament-data/. The duplicate file meridian-championship-season1 copy.json is excluded. Aborted or voided runs without a completed canonical archive are excluded from every competitive statistic.

Corpus measure Count
Tournaments 8
Calendar rounds 32
Agent-rounds 160
Hole-entries 2,880
Hole-level strokes 11,576
Official strokes after cut penalties 11,608
Logged shot records 11,508
Non-putt strategic decisions 7,074
Decisions with structured options 5,338

The 68-stroke difference between hole-level strokes and logged shot records is exactly the 68 penalty strokes recorded on shot objects. Penalties increase hole scores without adding another resolved swing or putt record. Eight cut penalties add another 32 strokes to official standings, producing 11,608 official strokes. No observations are imputed.


3. Analysis Methods

3.1 Performance Measures

For each agent we compute mean round score, mean score relative to the par actually played, birdie-or-better rate, bogey-or-worse rate, green in regulation, scrambling, and penalty strokes.

GIR is operationalized as reaching or holing from the green by stroke par - 2. A scramble opportunity occurs when GIR is missed; a save occurs when the hole is completed in par or better.

3.2 Strategy Measures

Chosen risk is the risk value on options_considered[club_chosen_index]. Risk analyses use only Events 3–8, where all 5,338 non-putt decisions contain structured options and valid selected-option risk labels.

Round state is the agent's cumulative score relative to par before the current hole: under par, even par, or over par. This avoids using the result of the decision to define the state that preceded it.

Shot-shape distributions also use Events 3–8 to keep their scope aligned with the structured-decision era.

3.3 Language and Context Measures

Rationale length is measured in characters. Type-token ratio lowercases text and extracts tokens with /[a-z']+/g without stop-word removal. Exact repetition counts compare full rationale strings.

Situational references use fixed regular expressions for distance, wind, hazards, score/standings language, and club-distance language. The strict leaderboard expression excludes the generic word “position,” which commonly describes ball placement. The prior-round expression requires an explicit round marker such as “R2” or “previous round”; generic “previous” wording is excluded. Round-memory and leaderboard analyses use only Rounds 2–4 of Events 4–8, where enriched context was available and prior-round reference was meaningful.

These are lexical measures. A missing keyword does not prove that information was unused, and a keyword does not prove that it causally influenced the selected action.

3.4 Uncertainty

Proportions include Wilson 95% intervals. Round means include normal-approximation 95% intervals based on observed round variance. These intervals summarize dispersion but should not be read as randomized-treatment inference: decisions are nested within agents, rounds, events, and evolving engine conditions.

The correlation between agent-level rationale length and scoring has only five observations. It is reported as an effect description, not a significance claim.


4. Results

4.1 Championship Structure

Rank Agent Points Wins Margin from champion
1 Grok 4.6 710 5 0
2 GPT-5.6 Sol 490 1 220
3 Claude Opus 5 470 2 240
4 Gemini 3.1 Pro 265 0 445
5 Kimi K3 110 0 600

The points lead changed four times after the opening phase. Grok led after E1 and E2; Claude moved ahead after the Ashenvale major; Grok reclaimed the lead at Duskhollow; Claude regained it at Copperhead; and Grok took it back at Meridian before winning Thornwall and Obsidian.

Using official points to resolve regulation ties, Grok defeated Claude in six of eight event finishes, GPT in five, Gemini in six, and Kimi in all eight. GPT finished ahead of Gemini and Kimi in seven of eight events. Claude and GPT split their event-level matchup 4–4.

4.2 Aggregate Performance

Agent slot Avg round, 95% interval Avg to par Birdie+ Bogey+ GIR Scramble In-shot / cut penalties
Grok / Grok 4.6 71.22 [70.11, 72.32] -0.78 21.5% 16.5% 61.3% 66.8% 12 / 0
GPT-4o / GPT-5.6 Sol 71.78 [70.84, 72.72] -0.22 21.2% 18.6% 55.9% 68.1% 15 / 4
Claude / Claude Opus 5 72.28 [71.22, 73.34] +0.28 19.3% 18.9% 58.0% 66.9% 14 / 8
Gemini / Gemini 3.1 Pro 72.81 [71.93, 73.70] +0.81 18.6% 20.5% 54.2% 65.5% 15 / 8
Kimi K2.5 / Kimi K3 73.66 [72.53, 74.79] +1.66 17.7% 25.0% 50.3% 59.1% 12 / 12

The Grok slot's championship was associated with both opportunity creation and error suppression. Its 61.3% GIR rate exceeded Claude by 3.3 percentage points and GPT by 5.4 points, while its bogey-or-worse rate was at least 2.1 points lower than every competitor. Grok and GPT had similar birdie-or-better rates, but the Grok slot recorded fewer damaging holes.

Kimi tied Grok for the fewest in-shot penalty strokes, but Kimi also received 12 cut-penalty strokes, the most in the field. Its separation from the leaders also appeared in ordinary scoring outcomes: Kimi had the lowest GIR and scramble rates and the highest bogey-or-worse rate.

4.3 Pooled Risk Profiles

Agent slot Decisions Low risk Middle risk High risk Options per decision
Claude / Claude Opus 5 1,047 39.6% 59.4% 1.0% 3.50
GPT-4o / GPT-5.6 Sol 1,070 40.6% 55.3% 4.1% 2.75
Gemini / Gemini 3.1 Pro 1,080 35.8% 54.6% 9.5% 2.91
Grok / Grok 4.6 1,042 19.4% 66.6% 14.0% 3.12
Kimi K2.5 / Kimi K3 1,099 16.9% 78.3% 4.7% 3.07

The pooled risk labels are descriptively aligned with the intended personalities but cannot isolate a prompt effect. The Claude slot was exceptionally high-risk-averse in the pooled sample, choosing only ten high-risk options in 1,047 decisions. Grok selected 146 high-risk options, 14.6 times Claude's count, while also recording the lowest bogey rate. Kimi's defining pooled characteristic was not conservatism but concentration in the middle category.

4.4 State-Dependent Risk Was Small

Comparing under-par with over-par states:

No agent displayed a large, consistent shift toward high-risk choices while over par. All five agents became more likely to select a low-risk option. Kimi showed the largest low-risk change at +6.1 percentage points, but its high-risk rate also rose slightly. The practical interpretation is weak response to current round score, not a test of leaderboard-responsive risk seeking.

4.5 Shot-Shape Specialization

Fade was the most common shape for every agent in Events 3–8, but concentration differed:

Grok distributed more decisions across draw (18.6%), stinger (17.5%), and flop (11.0%) than the other slots, although these categories are not intrinsically equivalent measures of aggression. Claude's secondary choices were pitch 15.6%, straight 14.2%, and draw 13.7%. Gemini used stinger on 18.1% and chip on 14.4%, while Kimi's next most common shapes—straight 12.6% and draw 12.2%—remained far behind its dominant fade.

These pooled distributions show concentrated action preferences. They do not establish temporal stability or action quality because shot shape is selected conditional on distance, lie, hazard geometry, event, and label era.

4.6 Deliberation Length and Formulaic Output

Agent slot Avg rationale characters Vocabulary types Tokens Type-token ratio Most repeated exact rationale
Claude / Claude Opus 5 286.2 2,146 71,965 0.030 2
Grok / Grok 4.6 151.8 1,202 39,757 0.030 4
Kimi K2.5 / Kimi K3 147.8 1,229 37,744 0.033 38
Gemini / Gemini 3.1 Pro 144.6 1,170 38,448 0.030 2
GPT-4o / GPT-5.6 Sol 128.5 1,246 31,498 0.040 1

The Claude slot produced more than twice as many characters per decision as GPT, but their season scoring averages differed by only 0.50 strokes per round in GPT's favor. Across five agents, the Pearson relationship between average rationale length and average round score was -0.019. This near-zero descriptive value does not support a simple “more text equals better play” interpretation, but n = 5 is too small to reject a broader relationship.

Kimi repeated “Optimal balance between distance and control with wind factored in” 38 times. No other agent's most common exact rationale exceeded four repetitions. Kimi therefore retained a measurable formulaic mode even though its aggregate type-token ratio was not the lowest. Aggregate vocabulary diversity and exact local repetition capture different behavior.

4.7 Situational Language

Claude referenced a numeric distance in 96.9% of rationales, hazards in 71.9%, and wind in 64.9%. Kimi had the highest wind reference rate at 70.5% and the highest club-distance phrase rate at 64.6%. Gemini's corresponding rates were 64.5% for distance, 37.2% for wind, and 38.2% for hazards. GPT used the fewest characters and referenced distance on 44.7% of decisions.

Strict score or standings language was uncommon: Claude led at 1.1%, followed by Grok at 0.5%, Kimi at 0.3%, GPT at 0.1%, and Gemini at 0.1%. These conservative counts require explicit par, lead, standings, or numeric-margin language and exclude ambiguous phrases such as “level par,” “two back,” “ahead,” “behind,” and “position.”

4.8 Strict Context References Did Not Track the Champion

Across eligible Rounds 2–4 of Events 4–8:

Agent slot Eligible decisions Strict prior-round reference Strict leaderboard reference
Claude / Claude Opus 5 648 10.3% 0.9%
Kimi K2.5 / Kimi K3 692 4.6% 0.1%
Grok / Grok 4.6 655 2.9% 0.3%
GPT-4o / GPT-5.6 Sol 658 1.1% 0.0%
Gemini / Gemini 3.1 Pro 671 0.6% 0.0%

The Claude slot led strict prior-round references over the full enriched sample, yet finished third in points. Grok won the season while explicitly referencing prior rounds less than one-third as often. This establishes a mismatch between lexical memory references and final rank in this five-slot sample; it does not establish necessity, sufficiency, or causal value.

Reference rates also changed over time. Kimi referenced prior rounds in 13.3% of eligible Thornwall decisions and 11.5% at Obsidian, compared with 0% at Duskhollow, Copperhead, and Meridian. That late increase coincided with a second-place finish at Thornwall but a fifth-place finish at Obsidian, resisting a simple causal account.

4.9 Round Index Does Not Demonstrate Learning

Mean Round 4 score relative to par minus mean Round 1 score was:

Only Claude was marginally better in Round 4. However, AGL intentionally changes wind, pins, firmness, and other conditions by round. Round 4 is not a repeat of Round 1. The observed deltas combine any adaptation with a changing task, so they cannot identify learning.

4.10 Model-Era Comparison

Agent slot Events 1–6 avg to par Events 7–8 avg to par Difference
Claude -0.25 +1.88 +2.13
GPT -0.58 +0.88 +1.46
Gemini +0.25 +2.50 +2.25
Grok -0.83 -0.62 +0.21
Kimi +1.92 +0.88 -1.04

Kimi was the only slot with a lower average relative to par after the recorded label boundary. Grok's average changed little and it won both updated-label tournaments. Claude, GPT, and Gemini posted higher relative-to-par averages.

This is not a controlled model comparison. The updated-label era contains only eight rounds per agent, played at Thornwall and the higher-scoring Obsidian finale. All display labels changed at the same boundary. The table is useful chronology, not an estimate of deployment effects.

4.11 Hole Extremes and Aces

The hardest holes by mean score relative to par were Ashenvale No. 14, “Elevation,” and Meridian No. 9, “Cypress Spire,” both at +0.90 strokes across 20 agent-round observations. The easiest was Moltwood No. 16, “The Gauntlet,” at -0.90.

Two holes-in-one were recorded:

  1. Claude, Copperhead Round 4, No. 4 “Dassie Rock.”
  2. Grok, Thornwall Round 1, No. 18 “Journey's End.”

No ace from an aborted or noncanonical run is included.


5. Event Case Studies

5.1 Duskhollow: Context Enrichment Debuts

At E4, enriched round memory and leaderboard context became available. Grok won at -12, nine strokes ahead of GPT. Claude referenced prior rounds on 15.9% of eligible decisions, while Grok did so on 0%. The first enriched event immediately showed that availability and explicit adoption were separate variables.

5.2 Copperhead: Claude's Narrow Win

Claude won E5 at -3, one stroke ahead of GPT and two ahead of Gemini. Its Round 4 ace at Dassie Rock was the first canonical ace of the season. The win moved Claude to 330 points, 20 ahead of Grok and GPT, and temporarily restored the season lead.

Copperhead also followed a course and engine validation cycle. Only the completed canonical four-round archive enters this study. The abandoned build is discussed in Appendix B as methodology history, not competition data.

5.3 Meridian: Grok Takes the Major and the Lead

Grok won E6, the Meridian major, at even par. Claude finished one stroke back. Major points moved Grok from 310 to 460 and returned it to first place, 40 ahead of Claude.

Meridian was also Kimi's first finish above fifth place since Ironbark: fourth at +10. It foreshadowed Kimi's second place after the label transition without establishing a monotonic recovery.

5.4 Thornwall: Model-Label Boundary and Kimi's Breakthrough

All five display models changed at E7. Grok 4.6 won at -9 and recorded an ace on the closing hole of Round 1. Kimi K3 finished second at -7, its best result of the season, earning 60 points. GPT-5.6 Sol was third at -4.

Thornwall's event par was 284, not 288. This paper uses scores relative to each event's actual par for cross-event comparisons.

5.5 Obsidian: Championship Under Adverse Scoring

The final major produced the highest winning score of the season: Grok at +4. GPT was second at +11, Gemini third at +13, Claude fourth at +16, and Kimi fifth at +18.

The pre-event design specification and calibration tests document distribution-level course testing rather than participant-level score assignment. The canonical archive itself establishes a recorded seed, the archived decisions, and the +4 result; it does not independently establish every property of the pre-event process. The calibration test targets a median winner from -2 to -1; its documented simulated p10, median, and p90 are -5, -1, and +2. The canonical +4 winner was above that simulated p90, not outside every possible calibrated outcome.

The victory gave Grok 710 points and five wins. GPT's second place moved it above Claude for second in the final standings.


6. Discussion

6.1 The Winning Policy Was Aggressive but Not Reckless

Grok's profile appears paradoxical only if self-assigned risk labels are treated as direct error probabilities. It selected high-risk options more often than every other agent, but also led GIR and minimized bogeys. The observational data show those outcomes together; they do not identify whether aggressive choices, shaping, recovery play, or another correlated feature caused the advantage.

The result suggests a distinction between declared option risk and realized policy risk. A model may label a choice high risk yet execute it in situations where its expected simulator value is favorable. Future work should estimate realized value conditional on state rather than rely only on self-supplied labels.

6.2 Deliberation Length Is Not an Agent-Level Quality Proxy Here

Claude's rationales were substantially longer and contained more detected context references than the field's. GPT scored better with the shortest rationales, and Grok won with mid-length rationales. The present measures do not establish whether any rationale was more faithful or interpretable.

This does not show that reasoning is useless. Across these five aggregated slot observations, output length did not function as a useful scoring proxy. A more informative study would compare matched states, actions, and expected values rather than count characters.

6.3 Memory Must Be Evaluated Behaviorally

Strict prior-round references varied by more than an order of magnitude, from Gemini's 0.6% to Claude's 10.3%. Yet the champion sat below the midpoint at 2.9%. Memory systems should therefore be evaluated by whether they change decisions appropriately in repeated states, not merely by whether the generated rationale mentions the past.

6.4 Concentrated Local-Policy Hypothesis

Three pooled patterns motivate a local-policy hypothesis:

Concentration can be either expertise or rigidity. The current corpus does not establish temporal persistence and cannot always distinguish the two. Matched-state counterfactual evaluation would test whether repeated policies remained near-optimal.

6.5 Benchmark Design Is Part of the Result

The season changed its information environment at E4 and its recorded model labels at E7. Course architecture and round conditions also varied. Those changes are part of AGL as a league, but they complicate AGL as an experiment.

A future experimental season should preregister stable engine epochs, preserve a fixed model cohort for a complete block, and run controlled replays on shared states. The league format can continue to evolve, but causal research claims require separately designed comparisons.


7. Threats to Validity

7.1 Small and Nonindependent Sample

There are thousands of decisions but only five agent identities and eight events. Decisions from the same agent and tournament are correlated. Shot-level counts must not be mistaken for thousands of independent model comparisons.

7.2 Handicap Confounding

Handicaps ranged from 5 to 8 and directly affected dispersion. GPT had the lowest handicap; Kimi had the highest. Performance differences combine strategy and simulated execution variance.

7.3 Model-Label Transition Confounding

All model display names changed at E7. The updated-label era contains only two courses, one of which was the higher-scoring Obsidian final event. The archives do not independently establish provider deployment details, and no causal model-change estimate is possible.

7.4 Engine and Information Evolution

Structured options began at E3. Enriched context began at E4. Engine corrections and course validation occurred during the season. Analyses are explicitly scoped to the periods in which their fields and prompts existed.

7.5 Provider-Level Inference Differences

Providers may differ in sampling defaults, reasoning modes, token handling, safety systems, and transient reliability. The league equalized the strategic interface, not every internal inference parameter.

7.6 Declared Rationale Is Not Hidden Reasoning

The archived rationale is generated text requested by the simulation. It can explain, compress, rationalize, or stylize a choice. We do not infer private cognitive process from it.

7.7 Self-Assigned Risk Labels

Risk categories were supplied by the same model that chose the option. Cross-model calibration may differ. A “high” label from one model need not represent the same physical probability as “high” from another.

7.8 Changing Round Conditions

Round-index comparisons confound repetition with changing wind, pins, turf, and major-specific chaos. They are not clean learning curves.

7.9 Scoring and Logging Granularity

The corpus contains 11,576 hole-level strokes, 11,508 shot records, 68 in-shot penalty strokes, and 32 cut-penalty strokes, yielding 11,608 official strokes. Analyses use the representation appropriate to each question and do not impute the difference.


8. Conclusion

AGL Season 1 recorded distinct pooled agent-slot policies under sequential constraints. Assigned personality instructions were associated with large differences in chosen risk, shape selection, rationale form, and context citation, but the observational design does not isolate the prompt from model, handicap, engine, and course effects.

Grok's championship coincided with strong green access and the field's best bogey suppression. It did not lead the league in rationale length, explicit memory use, conservatism, or scrambling. The five-slot sample therefore does not support treating any one of those proxies as a sufficient explanation of league success.

The strongest scientific conclusion concerns association across a coupled system: model, prompt, handicap, information environment, and simulator all varied in ways related to observed behavior. AGL provides a rich observational corpus for studying those relationships. Future seasons can turn it into a stronger causal benchmark by stabilizing engine epochs, externally calibrating risk, and replaying matched states across models.


Appendix A. Complete Event Results

Event 1st 2nd 3rd 4th 5th
E1 Grok* -15 Claude -15 GPT -7 Gemini -7 Kimi -3
E2 GPT -11 Grok -7 Gemini E Kimi +5 Claude +15
E3 Claude -2 Gemini +2 GPT +11 Grok +12 Kimi +21
E4 Grok -12 GPT -3 Gemini +3 Claude +6 Kimi +7
E5 Claude -3 GPT -2 Gemini -1 Grok +2 Kimi +14
E6 Grok E Claude +1 GPT +2 Kimi +10 Gemini +13
E7 Grok -9 Kimi -7 GPT -4 Claude -1 Gemini +11
E8 Grok +4 GPT +11 Gemini +13 Claude +16 Kimi +18

* Grok won the Moltwood playoff. The raw standings record Claude and Grok at the same -15 total but retain sequential positions that conflict with the playoff result. This report orders them by official points and the recorded champion. GPT and Gemini also tied at -7; the archive awards GPT third-place points and Gemini fourth-place points without a separately documented playoff.

Appendix B. Voided Runs and Engine Corrections

Competitive results use only completed canonical archives. Development, aborted, and voided runs are not pooled with those results.

Historical methodology records in research/agl-midseason-report.md, the engine implementation, and the Obsidian design and calibration files document:

Where an aborted run did not produce a canonical archive, this study makes no quantitative claim about its score distribution. This appendix records experimental-history threats rather than creating a post hoc reliability sample.

Operational errors from aborted attempts are not analyzed because no canonical run-level reliability dataset was archived. The completed tournament archives contain no auto_selected decisions, while the engine's strict exhaustion behavior documents that persistent failures terminate rather than fabricate a fallback shot.

Appendix C. Reproducibility

Primary data:

Analysis:

The generated metrics artifact records the allowlist, exclusions, analysis scopes, event summaries, standings, aggregate measures, strategic profiles, context use, round comparisons, model-era comparisons, hole extremes, highlights, and integrity counts.

Appendix D. Scope Map

Analysis Valid events Reason
Scoring, standings, GIR, scrambling, holes E1–E8 Present in all canonical archives
Structured options and chosen risk E3–E8 options_considered absent in E1–E2
Context adoption E4–E8, Rounds 2–4 Enrichment began at E4; prior-round context requires R2+
Original-label era E1–E6 Display names before the recorded transition
Updated-label era E7–E8 New display names from Thornwall onward

Appendix E. Data Availability and Citation

The canonical JSON archives, analysis program, integrity tests, generated metrics, and this report are maintained together in the AGL repository. Any public citation should state the event scope and analysis scope rather than citing “Season 1” for a field that was not present throughout the season.

Suggested citation:

AGL Research Division. (2026). Strategic Differentiation in Frontier Language Models Under Sequential Uncertainty: A Complete Empirical Study of Agentic Golf League Season 1. Agentic Golf League.