Skip to content

feat(narratives): update default Gemini model to gemini-3.8-flash - #503

Merged
nick-nlb merged 4 commits into
datacommonsorg:narratives-devfrom
nick-nlb:narr-model-5-model-bump
Oct 3, 2026
Merged

nick-nlb merged 4 commits into
datacommonsorg:narratives-devfrom
nick-nlb:narr-model-5-model-bump

Conversation

@nick-nlb

@nick-nlb nick-nlb commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Overview

This PR is the fifth stage of the Narratives model loop migration. It makes gemini-3.8-flash the default Gemini model and fixes the one incompatibility found when testing the agent against that model: the three lightweight structured calls (chart configuration, chart validation, and follow-up questions) requested thinking_level="minimal", which gemini-3.8-flash rejects with 400 INVALID_ARGUMENT: Thinking level MINIMAL is not supported for this model.

The rejection is easy to miss. Each of those calls is designed to fail quietly, so on the new model every turn would have completed normally with no charts and no follow-up questions, and nothing would have surfaced except a logged error. The unit tests mock Gemini, so CI could not catch it. The change was therefore verified against the live model as well as the unit suite.

Related Issues

This PR follows #494, #495, #496, and #501 in the Narratives model loop migration.

Changes Made

  • Lightweight thinking level: Chart configuration, chart validation, and follow-up questions now request the "low" thinking level through a shared constant. The unsupported "minimal" level was removed from the client's mapping so any caller requesting it falls back to "low".
  • Default model: Updated the default model from gemini-3-flash-preview to gemini-3.8-flash in config, schema defaults, example configs, and docstrings.
  • Tests: Added tests confirming the lightweight calls request "low" and that "minimal" normalizes to "low". Updated client test defaults to use the new model.

Behavior changes

  • New default model: Every Gemini call (tool loop, synthesis, chart configuration, chart validation, and follow-up questions) uses gemini-3.8-flash unless gemini.mcp_model is set.
  • Charts and follow-up questions work on the new model: Chart configuration, chart validation, and follow-up generation run at "low", the lowest level gemini-3.8-flash accepts. On the live model each finished in 1.5 to 3 seconds.
  • Optional chart fields arrive as null: gemini-3.8-flash returns explicit null for optional chart fields (parent_place, child_place_type, date) where the previous model omitted them. Every consumer already tolerates this: tile_chart.tsx sets each attribute only when the value is truthy, title falls back through ??, and variable_dcids defaults to [] in use_sse_chat.ts. No change was needed.

Testing Done

Verified the Python agent locally from narratives/agent/:

uv sync --frozen
uv lock --check
uv run ruff format --check .
uv run ruff check .
uv run mypy
uv run pytest -v

Ran the UI type-check, the full workspace test suite, and the production build from narratives/ with the Node version from .nvmrc:

nvm use
pnpm -C ui run lint
pnpm test
pnpm build

All 138 UI tests and 499 agent tests pass.

Manual verification against the live gemini-3.8-flash model:

  1. In narratives/agent/, ensure config.json sets gemini.api_key and does not override gemini.mcp_model (or sets it to gemini-3.8-flash), export MCP_SERVER_URL=https://api.datacommons.org/mcp, DATA_PLANE_URL=https://api.datacommons.org, DATA_PLANE_WEB_URL=https://datacommons.org, and DC_API_KEY, then run uv run narratives-agent-dev.
  2. In narratives/, run pnpm -C ui run dev and open http://localhost:3000.
  3. Ask "How has California's population changed since 2018?" and confirm that the turn streams reasoning thoughts and the synthesized answer, renders a line chart (get_chart_config and validate_data_response succeed at thinking_level="low" rather than failing with HTTP 400 on "minimal"), and emits follow-up questions on the follow_ups SSE event.
  • Unit tests passed
  • Integration tests passed [NOT APPLICABLE]
  • Manual verification
  • Any updated goldens or fixtures were reviewed and are intentional [NOT APPLICABLE]

Risk & Rollback

Low risk. Narratives has no active deployments. A deployment that sets gemini.mcp_model keeps its configured model. To roll back, revert the merge commit on narratives-dev.

Follow-ups

Checklist

  • I have read AGENTS.md and followed CODING_GUIDELINES.md, plus FRONTEND.md for UI changes.
  • I have run the app's lint, test, and build commands, as documented in that application's guide.
  • I have commented my code, particularly in hard-to-understand areas.
  • My changes generate no new warnings.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the default Gemini model from gemini-3-flash-preview to gemini-3.8-flash across the configuration, client, workflows, schemas, and tests. Additionally, because gemini-3.8-flash rejects the minimal thinking level, the minimal option has been removed, and a new LIGHTWEIGHT_THINKING_LEVEL constant set to low has been introduced and integrated into the chart configuration, validation, and follow-up workflows. New unit tests have been added to verify that these workflows correctly request the low thinking level. There are no review comments, and I have no additional feedback to provide.

@nick-nlb
nick-nlb marked this pull request as ready for review October 3, 2026 00:23
@nick-nlb
nick-nlb merged commit f3e68c0 into datacommonsorg:narratives-dev Oct 3, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants