The backend parsing engine for ArboratorGrew. A Flask + Celery REST API that wraps BertForDeprel for dependency parsing.
Quick_parsing (Flask frontend) ──HTTP──▶ arborator-parser (Flask + Celery backend) ──▶ BertForDeprel
- text tokenization - model management on disk
- file upload UI - Celery task queue (train / parse)
- proxies requests to backend - writes CoNLL-U, invokes BertForDeprel
- Quick_parsing is the user-facing frontend. It tokenizes text into CoNLL-U format and proxies all parsing/training requests via HTTP to this backend. It does not modify CoNLL-U content.
- arborator-parser (this repo) is the backend. It manages model files, runs BertForDeprel training and inference as Celery tasks, and returns parsed CoNLL-U.
- The frontend sends
X-Application-Models: quick_parserin request headers to select the flat model store used by Quick_parsing.
When a model's project_name starts with gloss_, the parse_sentences Celery task automatically performs a Gloss↔FORM swap:
- Pre-processing: For each token with
Gloss=Xin MISC, the original FORM is saved asOrigForm=in MISC and replaced with the gloss value. This lets BertForDeprel parse using English glosses instead of surface forms. - Post-processing: After parsing, the original FORM is restored from
OrigForm=and the gloss is placed back inGloss=in MISC. All predicted annotations (HEAD, DEPREL, UPOS, etc.) are preserved.
This is implemented in app/utils_gloss.py and called from app/models/celery_tasks.py.
- Clone server with submodules
# with ssh (git might ask you to set ssh key for the project)
git clone --recurse-submodules git@github.com:Arborator/arborator-parser.git
# or with http
git clone --recurse-submodules https://github.com/Arborator/arborator-parserCheck the doc on the official doc of redis
curl -fsSL https://packages.redis.io/gpg | sudo gpg --dearmor -o /usr/share/keyrings/redis-archive-keyring.gpg
- Install venv as you can (on the LISN machine, we need to set miniconda3 vens)
# if using conda/miniconda
conda create -n parser-venv python=3.8
conda activate parser-venv
pip install -r requirements.txt
# you might need to have a C compiler, if so, `
sudo apt install build-essential
# and this dependency for uwsgi
sudo apt-get install libxcrypt-dev
# if using python-venv
python3.8 -m venv parser-venvTODO : Document this part
# if the dir does not exist
mkdir -p ~/.config/systemd/user/
# then link the service and enable it
ln arborator-parser.service /etc/systemd/system
systemctl enable arborator-parser.service
systemctl start arborator-parser.service- Create service for user Celery app (can't use system-wise service as arboratorgrew is not root). The template of the service is at the root of this repo.
ln arborator-parser-celery.service ~/.config/systemd/user/
systemctl enable arborator-parser-celery.service
systemctl start arborator-parser-celery.serviceAdd the nginx server block conf file. Again, it can be found at the root of this repo.
# access logs :
sudo tail -f /var/log/nginx/access.log
# and error logs :
sudo tail -f /var/log/nginx/error.logAll service logging :
journalctl -f
And our services logs :
# for flask app
systemctl status arborator-parser.service
# and for the celery workers (parse queue / train queue)
systemctl status arborator-parser-celery.service
systemctl status arborator-parser-celery-train.serviceLogs of arborator-parser (from root of this repo)
# live tail of the flask logs
tail -f ./logs/arborator-parser.log
# live tail of celery service logs, added to the previous 100 lines
journalctl --unit=arborator-parser-celery.service -f -n 100- PORT : 8002
- Path : /home/arboratorgrew/arborator-parser/wsgi.py
- PATH_MODELS : /home/arboratorgrew/arborator-parser_models/
- specific config files : arborator-parser-celery.service ; arborator-parser-celery.service ; arborator-parser.ini ; arborator-parser.nginx.conf ;
- socket : arborator-parser.sock
/!\ Don't install the dev server on the same user as the prod server. Indeed, the dev and prod version of celery will collide.
- PORT : 8001
- Path : /home/arboratorgrew/arborator-parser_dev/wsgi.py
- PATH_MODELS : /home/arboratorgrew/arborator-parser_models_dev/
- specific config files : arborator-parser-celery.service ; arborator-parser-celery_dev.service ; arborator-parser_dev.ini ; arborator-parser_dev.nginx.conf ;
- socket : arborator-parser_dev.sock
- It's using the same reddis port as the production server, we want to change that to have each of them using their own instance of reddis.
First check all the logs mentioned above
Then, check healthyness : On the server, do the following curl :
curl localhost:8002/parser/healthy
# or
curl localhost:8012/parser/healthyIf port tunneling to server, you can access to the doc by going on the URL : https://127.0.0.1:8088/parser/doc
To be sure that the server is always running, we added systemctl timer
it should check the server healthy endpoint, and restart the arborator-parser.service and arborator-parser-celery.service
Make sure the script is executable with
chmod +x /path/to/your/script.sh
template of the service is in this repo under the name check_server.service. It should point to the correct check_server.sh script
template of the timer is in this repo under the name check_server.service. It should point to the correct check_server.service
systemctl --user daemon-reload
systemctl --user enable --now check_server.timer
As we use user services, when the user session scope terminate (can be minutes, hours or days after the user leave the session), the system will stop all of the services of this user.
To prevent this, we set
sudo loginctl enable-linger arboratorgrew
systemctl --user status check_server.service
systemctl --user status check_server.timer
journalctl --unit check_server.service
journalctl --unit check_server.timer
The box has a single GPU shared with other services, so a parse/train job can
find the GPU short of memory. GET /parser/status reports, in one JSON:
How it works:
- Tasks are routed to two queues (
task_routesinapp/__init__.py):parseandtrain, each served by its own worker with--concurrency=1(arborator-parser-celery.service→parse@host,arborator-parser-celery-train.service→train@host). So a multi-hour training never blocks parsing, at most one job of each type touches the GPU, and the queue position per type is exact. The parse worker also drains the defaultceleryqueue for messages enqueued by an older server version. - Each task writes a small sidecar
logs/tasks/<task_id>.jsonand captures BertForDeprel's stdout inlogs/tasks/<task_id>.log. BertForDeprel printsTraining: 12.50% complete. 173.42 seconds in epoch (85.00 sents/sec)style lines; the ETA is computed from those, so nobody else on the GPU has to report anything. Training gives two ETAs because of early stopping (patience):eta_min_s(stops as soon as patience allows) andeta_max_s(runs tomax_epoch). - Before launching, a task waits (parse: 3 min, train: 15 min, see
GPU_WAIT_MAX_S_*) untilmemory_free_mib >= gpu_requirements_mib[type], showingphase: waiting_for_gpumeanwhile; on timeout it fails with agpu_busymessage instead of crashing with a CUDA OOM. - While a task runs, its peak GPU memory is sampled with nvidia-smi and stored in
the sidecar. After 3 successful runs of a type, the requirement becomes
max(last 20 peaks) * 1.2, so the thresholds calibrate themselves. Until then the defaults (GPU_MIB_PARSE,GPU_MIB_TRAINin.flaskenv) apply. Measured on the A6000 with xlm-roberta-large: predict, batch 8 → 2.6 GB peak, ~10 s model load + ~40 sents/s. POST /parser/models/train/statusand/parse/statusnow also returndata.progress(same object as inrunning[].progress, or{"phase": "queued", "queue_position": n, "estimated_wait_s": s}), so the existing polling loop in ArboratorGrew / Quick_parsing gets the ETA for free.
Restarting a worker kills the task it is running or waiting on (a task in
waiting_for_gpu is inside the worker too). Before restarting, check that
running and queued are empty in curl localhost:8002/parser/status — a queue
length of 0 and no run.py process is not enough. Then:
sudo -u arboratorgrew git -C /home/arboratorgrew/arborator-parser pull origin main # never `sudo git pull` (root-owned .git objects)
sudo cp /home/arboratorgrew/arborator-parser/arborator-parser-celery*.service /etc/systemd/system/ && sudo systemctl daemon-reload
sudo systemctl restart arborator-parser.service arborator-parser-celery.service arborator-parser-celery-train.serviceEnv vars (.flaskenv): GPU_INDEX (0), GPU_MIB_PARSE (3500), GPU_MIB_TRAIN
(16000), GPU_WAIT_MAX_S_PARSE (180), GPU_WAIT_MAX_S_TRAIN (900), PATH_TASKS
(<repo>/logs/tasks), NVIDIA_SMI
(/usr/bin/nvidia-smi; the uwsgi unit's PATH does not include /usr/bin).
{ "healthy": true, "workers": {"parse": true, "train": true}, // the per-queue Celery workers answer a ping "gpu": { // straight from nvidia-smi "name": "NVIDIA RTX A6000", "memory_total_mib": 49140, "memory_used_mib": 776, "memory_free_mib": 47765, "utilization_pct": 0, "used_by_parser_mib": 0, "used_by_others_mib": 776, // "others" = services outside the parser; no ETA for those "processes": [{"pid": 2214, "name": ".../python", "used_mib": 766}] }, "gpu_requirements_mib": {"parse": 3500, "train": 16000}, // free memory needed to launch "availability": { // the answer to "can I start, and if not, when?" "parse": {"can_start_now": true, "blocked_by": null, "estimated_wait_s": 0, "memory_ok_after_queue": true}, "train": {"can_start_now": false, "blocked_by": "queue", "estimated_wait_s": 4145, "memory_ok_after_queue": true} }, // blocked_by: null | "queue" | "gpu_memory" | "worker_offline" "can_start_now": {"parse": true, "train": false}, // shortcut of availability[*].can_start_now "running": [{"task_id": "...", "type": "train", "project_name": "...", "started_at": 1758020000, "progress": {"phase": "training", "epoch": 3, "max_epoch": 10, "percent": 31.2, "eta_min_s": 600, "eta_max_s": 4100, "epochs_without_improvement": 1}}], "queued": [{"task_id": "...", "type": "parse", "project_name": "...", "position": 1, "estimated_duration_s": 45}], "estimated_wait_s": {"parse": 0, "train": 4145}, // per type: queue drain time for a job submitted now (null if unknown) "typical_duration_s": {"parse": 45, "train": 3600} }