Skip to content

Merge development into production - #483

Merged
rbruhn merged 3 commits into
mainfrom
development
Oct 4, 2026
Merged

rbruhn merged 3 commits into
mainfrom
development

Conversation

@rbruhn

@rbruhn rbruhn commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Summary

  • promote the current development branch to production

Included

Deployment

Before merging, stop the Panoptes listener on production:

sudo supervisorctl stop biospex:panoptes-pusher

The deploy then runs the migration followed by the cleanup (update_queries_operation = pusher-transcription-duplicates); expect a few minutes on the shared MongoDB server.

After the deploy, start it again and confirm only one listener runs:

sudo supervisorctl start biospex:panoptes-pusher
sudo supervisorctl status | grep -i panoptes

Classifications missed while it was stopped are backfilled by the reconcile chain (ZooniversePusherJob). No events are running.

Follow-up: set update_queries_operation back to ''.

Validation

  • development deployment completed on a fresh copy of production pusher_transcriptions with production's indexes:
    • 3,987,992 → 3,586,761 documents, one per classification_id (76,843 duplicated ids, 401,231 extra documents removed)
    • classification_id_1 is UNIQUE; the 711-copy classification 703618983 kept its earliest copy
    • event_transcriptions migration ran and its unique index exists

rbruhn added 3 commits October 4, 2026 11:46
Add a unique index on event_transcriptions (classification_id, event_id, team_id, user_id) and record rows with createOrFirst(), so concurrent listeners or retried jobs cannot double count a transcription and a duplicate no longer fails the job. Production has no duplicate rows, so the index can be added by a regular migration.
Add the pusher-transcription-duplicates update operation and run it on deploy. For each classification_id it keeps the earliest document and deletes the rest, then rebuilds the classification_id index as unique so PusherTranscriptionJob's existing E11000 handling stops further copies from duplicate listeners or retried jobs. The cleanup repeats if new duplicates arrive before the index is built, and a re-run is a no-op.

Closes #481
Remove duplicate transcriptions and enforce unique classification ids
@rbruhn
rbruhn merged commit fc69384 into main Oct 4, 2026
4 checks passed

This branch was successfully deployed

1 active deployment
development — 4c9799ea Deployed Oct 4, 2026 by rbruhn via build-and-deploy-development #281
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Remove duplicate Notes From Nature transcriptions and enforce unique classification ids

1 participant