Skip to content
 
 

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

HumorLens

This repository contains the prompts and minimal running scripts for HumorLens, a knowledge-grounded and rhetoric-aware framework for interpreting offensiveness in Chinese dark humor. It includes the standard multi-label baseline prompt, the rhetoric-aware agent prompts, and lightweight scripts showing how to call the prompts. It does not contain API keys or the full dataset.

Data Access

The full HellHumor-CN dataset contains sensitive offensive dark-humor content. For responsible release, the data files are not included in this public repository. Researchers who need the dataset for academic use should contact the authors and request access. After approval, place the dataset file at data/hellhumor_cn.jsonl or pass its path through the --input argument.

Contents

  • prompts/baseline_multilabel_zh.txt: standard multi-label prompt for offensiveness, topic, target, and explanation.
  • prompts/agent_prompts_zh.json: prompts used by the agent pipeline, including entity extraction, knowledge summarization, tragedy-trivialization analysis, irony analysis, pun analysis, and final aggregation.
  • src/run_baseline.py: minimal script for running the baseline prompt through OpenRouter.
  • src/run_agent.py: minimal script for running the agent-style prompt pipeline through OpenRouter. It can automatically retrieve entity knowledge after entity extraction.
  • src/retrieve_knowledge.py: optional standalone retriever using Serper, Wikipedia, and Baidu Baike.
  • data/README.md: dataset access note.

Setup

cd HumorLens
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

API Key

No API key is stored in this repository. Either set an environment variable:

export OPENROUTER_API_KEY="YOUR_KEY_HERE"

or leave it unset; the scripts will ask you to enter the key manually at runtime.

Run Baseline Prompt

python src/run_baseline.py \
  --input data/hellhumor_cn.jsonl \
  --model google/gemini-2.5-flash \
  --output outputs/baseline_predictions.jsonl

Retrieve Entity Knowledge

The full system retrieves background knowledge for detected entities. src/run_agent.py performs this step automatically after the entity-extraction prompt, then passes the retrieved evidence through the knowledge-summary prompt. Set SERPER_API_KEY, or leave it unset and enter it manually at runtime.

You can also run the retriever as a standalone preprocessing step:

export SERPER_API_KEY="YOUR_SERPER_KEY_HERE"
python src/retrieve_knowledge.py \
  --input data/hellhumor_cn.jsonl \
  --output outputs/hellhumor_cn_with_retrieval.jsonl

The output JSONL adds entity_search_results and retrieved_knowledge, which can then be passed to the agent script. If this field is already present, src/run_agent.py uses it directly instead of searching again.

Run Agent Prompt Pipeline

python src/run_agent.py \
  --input data/hellhumor_cn.jsonl \
  --model google/gemini-2.5-flash \
  --output outputs/agent_predictions.jsonl

The agent flow is: entity extraction -> search retrieval -> knowledge summarization -> tragedy-trivialization analysis -> irony analysis -> pun analysis -> final judgment. If you want to ablate retrieval, pass --no_retrieval.

Output Format

Both scripts write JSONL files. Each row contains the original text, prompt(s), raw model output, and a best-effort parsed JSON object.

Expected model output format:

{
  "offensive": 0,
  "topic": [],
  "target": null,
  "explanation": "..."
}

For offensive examples, topic is a multi-label array with IDs from 1 to 8.

Topic IDs

  1. Death & violence
  2. Disease, disability & body
  3. Sex, romance & reproduction
  4. Identity, groups & discrimination
  5. Disasters & large-scale suffering
  6. Family, childhood & orphans
  7. Religion, sacredness & philosophy
  8. Other

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages