Gameplay telemetry, read as evidence

A reading a game designer can check.

A game studio brings a GameAnalytics export. An agent reports what the play sessions did, and which session to look at. The game designer then opens that part in the game and plays it, to decide what to change.

This page demonstrates that process. Each claim is limited to what the cited studies support. The Super Mario section is a simulated case, not telemetry from a released game. A real export is read the same way. The reading describes what these sessions did.

4readings for a game studio
12papers this project answers
6skills in the repository
1mapping, confirmed once

What this project is for

A studio can use these readings in more than one place. Two of those uses are below. Both start from the same export and the same mapping. Both run exploration, risk-taking, and experimentation first, then play style last.

Use 1 · After the sessions

The designer decides what to change.

The designer brings a GameAnalytics export and confirms what each design id means. She runs the three readings on that pair, then gives those three reports to the play-style skill. Play style adds no score of its own. It places the three readings beside the finishes, the failures, and what was spent, and it names the session to open. That session is the place she looks when she decides how to improve the game.

Use 2 · During play

A game that already calls a model.

Some games already send a prompt to a language model while the player is playing. The game asks what she wants to do, or she chooses: enter the cave, or jump into the pipe. That choice is the prompt. The game already holds the log, and it can run these four skills on it. The agent runs the three readings, then play style. Suppose the cave is the harder option and the pipe is the safer one, and the risk-taking report shows that these sessions rarely took the harder option when a safer one was available. The agent can suggest the pipe. The suggestion follows that count. The report does not name a kind of player.

The research, and what we do about it

Twelve papers. Each paper leaves a problem open. This project answers that problem by limiting what a reading is allowed to claim. The full argument is in our approach document. The shape of the export, and the rule that a mapping must be confirmed, are in our log contract.

Experience

The problem

A metric is not an experience

Yannakakis, Spronck, Loiacono, and André show that gameplay metrics observe experience only indirectly. Pedersen, Togelius, and Yannakakis could predict challenge and frustration more readily than fun, and the same signal, such as a death, raised those states differently.

Our answer

The report keeps observation, measurement, inference, and interpretation apart. It names the game conditions the count depends on. It does not predict fun, frustration, or arousal.

Transfer

The problem

A model from one game does not travel by itself

N. Shaker, M. Shaker, and Abou-Zleikha found that features which did not mean the same thing in two games did not survive the translation. Melhart, Liapis, and Yannakakis predicted arousal change inside one genre. Time dominated. They do not claim the result across genres.

Our answer

A reading moves to another game only after the studio confirms that the event ids share a meaning.

Sequence

The problem

Totals hide order

Chen, Seif El-Nasr, Canossa, Badler, Tignor, and Colvin argue that aggregate counts discard the sequence of choices. Their sequences tracked expertise more than personality. This citation is the abstract. Bakkes, Spronck, and van Lankveld show that a button-level model fails when the same style is played with different controls.

Our answer

Exploration, risk-taking, and experimentation read sequences and choices. Play style sets those readings beside progression and resource. None of them emit a personality score.

Exploration

The problem

Exploration depends on the goal in the session

Gómez-Maureira, Kniestedt, van Duijn, Rieffe, and Plaat found that an explicit goal and being paid reduced spatial exploration, and that a curiosity trait did not predict who explored. Acevedo, Choi, Liu, Kao, and Mousas found that a coin-collection task lowered the share of the map players visited.

Our answer

The exploration reading counts optional-area events while a mapped goal or reward is in the session. A missing goal or reward is missing context. It is not a low score.

Risk

The problem

A general risk questionnaire is a weak fit

Lyu, Zhao, Zhang, Chen, Zhou, and Zhu paired Dota 2 matches with a general risk-propensity questionnaire. The best model explained about 17% of the questionnaire.

Our answer

Risk-taking records a harder option taken while a safer one was available in the log. The later win or failure stays beside that choice. A failure on its own is not the score.

Mechanics

The problem

What to measure is decided with the game

Seif El-Nasr describes every action as what, where, when, and who, and treats useful features as the ones tied to the core mechanics, chosen with the people who designed the game. A designer can check over-used and under-used areas, features not used as intended, and where progression sticks.

Our answer

That choice is the mapping. The agent asks only about design events the schema does not already explain.

Language

The problem

The designer gets sentences tied to events

Rubio-Manzano and Triviño turn gameplay numbers into sentences through rules the designer defined, so a session is more than a scoreboard.

Our answer

Every sentence in the report traces to a measurement. The report does not invent a label the events do not support.

Motivation

The problem

Why is not in the stream

Voitovich states the gap between telemetry, which records what players do, and motivation models, which talk about why. The paper offers a research agenda. This citation is the abstract.

Our answer

The report does not assign a motivation type.

Six skills

Four skills produce the readings a studio asks for. Two skills are for people changing this repository. A studio does not install packages to use the four.

01 · Exploration

Did players visit optional places while a goal was available?

It counts optional-area events, and how many sessions contain them, under the goal or reward the mapping names.

Example. Three sessions finish the level. Two also visit a cave or a grove. Score: 3 optional-area events in 2 of 3 sessions. The session that only finishes the level stays in the count.

02 · Risk-taking

Did players pick the harder option while a safer one was also there?

Among sessions that show the safer alternative, it counts how many also took the harder option. The later win or failure stays beside that count.

Example. One session takes the side door and fails. One takes the side door and the boss, then finishes. Score: 1 of the 2 sessions that showed the safer alternative also took the harder option.

03 · Experimentation

Did players try a different option on a later attempt?

It counts sessions where the mapped option changed, and it keeps the order. A later win does not prove the change worked. No paper in this set studies experimentation by that name. The reading follows the sequence argument above.

Example. Sword, then fail, then bow. Sword only, then complete. Bow, then fail, then staff. Score: the equipped option changed in 2 of 3 sessions.

04 · Play style

Which session should the designer open?

It repeats the three scores, then counts finishes, failures, and what was spent. It names the session that failed and what that session spent. It does not produce a style score, and it does not re-read design events.

Example. Two sessions finish level 1. Session 2 fails level 1 and spends 10 gold. That is the session to open.

How a studio uses it

No package install. The skills are the files under .agents/skills/. Open this folder as the agent's workspace.

1

Bring the export

Bring one file that holds the game's events. Each event inside that file is its own JSON object. Our log contract is the document in this project that describes that shape.

2

Confirm the mapping

A mapping is a file that says what each design event means for these readings: an optional place, a harder option, a change of power, or an event to leave out. The agent helps you confirm it. It recommends a meaning and asks short questions until you agree. Save that file beside the export.

3

Ask for three readings

Exploration, risk-taking, and experimentation, in any order. Each one reads the same export and the same mapping.

4

Ask for play style last

Play style is not another score. It is a summary. It places the three readings beside the finishes, the failures, and what players spent in the game, such as coins or lives. It then names the session to open. Ask for that report last. Give it the three readings and the same export. If one reading is missing, it stops and does not invent it.

What a designer does with a report

A higher count means the behavior showed up in more sessions. You compare that count to the behavior you built the game for. The same 2 of 3 can be the outcome you hoped for, or a sign the level is losing people. The skill reports the count. You decide the change.

Exploration · 2 of 3 sessions visited an optional place

You look at those places against the goal you set.

You built the cave and the grove so players would find them. Then 2 of 3 means most sessions reached them. You built them as rare detours. Then the same 2 of 3 means the main path is losing people. You walk the level with that count in mind and decide whether the optional places, or the goal, need a change.

Risk-taking · 1 of 2 sessions took the harder door

You play both doors.

The score is only the sessions where the safe door was also available. You built the hard door as an optional challenge. Then 1 of 2 is players using it. You hoped players would keep the safe door. Then 1 of 2 is players walking into the hard one. The failure stays a fact about the level. It is not, by itself, the risk.

Experimentation · the weapon changed in 2 of 3 sessions

You read the order, then you try that order.

You built several weapons so players would try them. Then 2 of 3 is that plan showing up. You built one weapon as the intended one. Then 2 of 3 is players leaving it. A later finish does not prove the new weapon caused the win. You play the sequence the report lists and see what the change did in the fight.

Play style · session 2 failed and spent 10 gold

You open that session.

The other three readings already said how often each behavior showed up. This one adds which session finished, which session failed, and what that session spent. You open session 2. You look up session 2 in the other readings. The behavior they recorded for it is the path you play: the safe door, the hard door, a weapon change, or an optional place. That path is the place to change, because that is where a session lost and spent gold. If those readings never name session 2, you still open the failed level, and you do not guess which behavior went with the loss.

Where the report stops

The agent names the count and the session. You play that part of the game and choose the change. The report does not rank a larger count as a better game, and it does not name a kind of player.

A simulated case: Super Mario Bros.

This section records one agent session on a simulated GameAnalytics export. It is not telemetry from a released game. The designer reads the log, confirms the mapping, and then runs each skill. Steps 3 to 6 summarize those reports. They do not replace the reports.

Step 1

Read the log before mapping any design id

The export contains two players and four sessions, on build 1.0.0. The lines below state what happened. They are not scores. The designer reads them before confirming a mapping, because an id cannot be classified until the session that produced it is understood.

  • Player 1, session 1 lasts 128 seconds on world 1-1. The player hits a question block, stomps a goomba, collects a mushroom, and enters the coin-room pipe. A pit costs a life, and the level fails. On a second attempt the player collects another mushroom, stomps a koopa, reaches the flagpole, and completes the level with a score of 4800.
  • Player 1, session 2 lasts 73 seconds on world 1-2. The player enters the warp-room pipe and takes the warp to world 4. World 1-2 is recorded as a failure. World 4-1 then starts. Lakitu hits the player, a life is lost, and that level fails.
  • Player 1, session 3 lasts 64 seconds, again on world 1-2. The player collects a fire flower, hits a piranha plant with a fireball, leaves by the goal pipe, reaches the flagpole, and completes the level with a score of 3200. This session does not enter the warp room.
  • Player 2, session 1 lasts 74 seconds on the overworld of world 1-1. The player hits a question block, is hit by a goomba, and loses a life. On a second attempt the player stomps a goomba, reaches the flagpole, and completes the level with a score of 2100. This session does not enter the coin-room pipe.

Step 2

Confirm one mapping for the three readings

Completions, failures, coins, and lives already have a meaning in the GameAnalytics schema, so they are not mapped. For each design id, the designer states whether it is an optional place, a harder option, a power that can change, or none of those. An entry marked meaning: ignore remains in the export for the GameAnalytics dashboard and is omitted from these four readings. One mapping file serves all three readings.

  • pipe:enter:bonus is the pipe into the coin room, while world 1-1 can still be finished. It is exploration, with the goal Complete:world1:level1.
  • pipe:enter:warpzone is the pipe that leaves the main path of world 1-2. It is exploration, with the goal Complete:world1:level2.
  • area:warpzone is that room, and it uses the same goal.
  • powerup:* is the equipped power, either a mushroom or a fire flower. It is experimentation.
  • warp:use:* is the skip to world 4. It is risk-taking. The safer way out of world 1-2 is pipe:exit:goal.
  • block:* is a question-block hit. The coin is already a resource event, so this id is ignored.
  • enemy:* is a stomp, a fireball, or a hit. It is ignored.
  • hazard:* is the pit. The lost life is already a resource sink, so this id is ignored.
  • goal:* is the flagpole. Finishing is already a progression event, so this id is ignored.
  • pipe:exit:* is leaving a pipe. It is ignored as exploration. pipe:exit:goal remains the safer exit named on the risk-taking line.
  • area:overworld is the main path, so it is ignored.
  • area:underground is world 1-2 itself, not an optional room, so it is ignored.

Step 3

Run the exploration skill

The text below is a summary, not the full report. It is the result of running x-exploration-reviewer on the log and the mapping above. The skill counts how many sessions that contain the named goal also contain the optional place.

One of the two sessions that complete world 1-1 also enters the coin room.

The warp room is not included in that score, because it was entered in a session that does not record a completion of world 1-2. Player 2 completes world 1-1 without entering the coin-room pipe.

Step 4

Run the risk-taking skill

The text below is a summary, not the full report. It is the result of running x-risk-taking-reviewer on the same log and the same mapping. The skill counts how many sessions that show the safer alternative also take the harder option.

The score is withheld because the warp and the goal pipe occur in different sessions.

The warp occurs in player 1, session 2. The goal pipe occurs in player 1, session 3. No session contains both, and a single session is not sufficient for a score. The failure on world 4-1 remains a fact about that level. It is not the risk-taking score.

Step 5

Run the experimentation skill

The text below is a summary, not the full report. It is the result of running x-experimentation-reviewer on the same log and the same mapping. The skill counts how many sessions change the equipped power within that session.

The equipped power changed in none of the four sessions.

The mushroom is collected twice, both times in player 1, session 1. The fire flower is collected once, in player 1, session 3. A different power in a later session does not raise the score. The completion of world 1-2 does not show that the fire flower caused the finish.

Step 6

Run the play-style skill last

The text below is a summary, not the full report. It is the result of running x-play-style-reviewer after the three readings above. This skill does not produce a score of its own. It places those readings beside the completions, the failures, and what was spent, and it names the session the designer should open.

The designer should open player 1, session 2. There is no single play-style score.

That session takes the warp, does not take the goal pipe, fails on world 4-1, and spends one life, in 73 seconds. Across the export there are 3 completions, 4 failures, 40 coins gained, and 3 lives spent. The pit in session 1 is a separate life, in the session that did enter the coin room, and it is not part of the warp.

Step 7

What the designer and the studio take from the reports

The counts are read against the behavior the levels were built to produce. The coin room was intended to be found. The warp was intended as a choice beside the goal pipe. Both powers were intended to be tried within one attempt.

What was wrong

  • World 1-1 can be completed without the coin room. One of the two completions never enters it.
  • World 1-2 never shows the warp and the goal pipe in the same session. The session that takes the warp then fails on world 4-1.
  • No session tries both powers. The fire flower appears only on a later visit.

What you change

  • Make the coin-room pipe visible on the way to the flagpole of world 1-1.
  • Place the warp and the goal pipe so that one session can encounter both, and continue to log both.
  • Offer the fire flower on the first visit to world 1-2.
  • Play player 1, session 2, before changing world 4-1.

What you would have missed

  • The longest session ends with a score of 4800. The export alone points there. The play-style report points to the 73-second session that takes the warp and then fails.
  • A completion does not show that the optional room was entered. Player 2 completes world 1-1 and never enters the coin room.
  • A death after a warp can be mistaken for risk-taking. The safer exit is absent from that session, so the score is withheld.
  • A mushroom followed later by a fire flower can be mistaken for experimentation within a session. No session changes the equipped power, so the score is 0 of 4.

This export contains no business event, so it does not show a purchase or a change in revenue. The four reports separate the three findings above. The export alone does not make that separation.

The twelve references

  1. G. N. Yannakakis, P. Spronck, D. Loiacono, and E. André, “Player Modeling,” in Artificial and Computational Intelligence in Games, Dagstuhl Follow-Ups, 2013.
  2. C. Pedersen, J. Togelius, and G. N. Yannakakis, “Modeling Player Experience for Content Creation,” IEEE Transactions on Computational Intelligence and AI in Games, vol. 2, no. 1, pp. 54–67, 2010. doi.org/10.1109/TCIAIG.2010.2043950
  3. N. Shaker, M. Shaker, and M. Abou-Zleikha, “Towards Generic Models of Player Experience,” in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, vol. 11, no. 1, 2015. doi.org/10.1609/aiide.v11i1.12806
  4. D. Melhart, A. Liapis, and G. N. Yannakakis, “Towards General Models of Player Experience: A Study Within Genres,” in IEEE Conference on Games, 2021. arxiv.org/abs/2110.00978
  5. Z. Chen, M. Seif El-Nasr, A. Canossa, J. Badler, S. Tignor, and R. Colvin, “Modeling Individual Differences through Frequent Pattern Mining on Role-Playing Game Actions,” in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, vol. 11, no. 5, pp. 2–7, 2015. Abstract. doi.org/10.1609/aiide.v11i5.12847
  6. S. C. J. Bakkes, P. H. M. Spronck, and G. van Lankveld, “Player behavioural modelling for video games,” Entertainment Computing, vol. 3, no. 3, pp. 71–79, 2012. doi.org/10.1016/j.entcom.2011.12.001
  7. M. A. Gómez-Maureira, I. Kniestedt, M. van Duijn, C. Rieffe, and A. Plaat, “Level Design Patterns That Invoke Curiosity-Driven Exploration,” Proceedings of the ACM on Human-Computer Interaction, vol. 5, CHI PLAY, article 271, 2021. doi.org/10.1145/3474698
  8. P. Acevedo, M. Choi, H. Liu, D. Kao, and C. Mousas, “Game Level Design to Evoke Spatial Exploration: The Influence of a Secondary Task,” in Companion Proceedings of the Annual Symposium on Computer-Human Interaction in Play, 2024. doi.org/10.1145/3665463.3678811
  9. S. Lyu, N. Zhao, Y. Zhang, W. Chen, H. Zhou, and T. Zhu, “Predicting Risk Propensity Through Player Behavior in DOTA 2,” Frontiers in Psychology, vol. 13, 2022. doi.org/10.3389/fpsyg.2022.827008
  10. M. Seif El-Nasr, “Intro to User Analytics,” Game Developer, 2013. gamedeveloper.com
  11. C. Rubio-Manzano and G. Triviño, “Automatic Linguistic Feedback in Computer Games,” in Proceedings of the 2015 Conference of the International Fuzzy Systems Association and the European Society for Fuzzy Logic and Technology, 2015. doi.org/10.2991/ifsa-eusflat-15.2015.69
  12. I. Voitovich, “Integrating Telemetry with Player Motivation Models,” in Conference Proceedings of DiGRA Central Asia, 2026. Abstract. dl.digra.org