Official implementation and evaluation artifact for ClawRec: A Claw-Native Recommender System.
ClawRec is an agentic recommender system that turns authorized cross-platform browser behavior into an evidence-linked user state. It plans retrieval by source role, acquires candidates through a visible browser session, curates a complementary recommendation slate, renders a local generative interface, and feeds explicit interaction signals back into the user state.
flowchart LR
A["Authorized browser behavior"] --> B["Unified events"]
B --> C["Evidence-linked user state"]
C --> D["Content planning"]
D --> E["Cross-source acquisition"]
E --> F["Marginal curation"]
F --> G["Generative interface"]
G --> H["Explicit feedback"]
H --> C
Each stage is implemented as an independent agent Skill with a narrow input and output contract:
| Stage | Skill | Primary artifact |
|---|---|---|
| Lifecycle routing | clawrec-using-clawrec |
Selects the required stage sequence |
| Behavior observation | clawrec-observing-browser-behavior |
UnifiedEvent JSONL |
| User-state synthesis | clawrec-synthesizing-user-state |
USER_STATE.md |
| Retrieval planning | clawrec-making-content-plan |
PLAN.md |
| Search orchestration | clawrec-orchestrating-search-execution |
Search task results |
| Source execution | clawrec-executing-search-instruction |
Verified source records |
| Recommendation curation | clawrec-curating-display-data |
display-data.json |
| Interface rendering | clawrec-rendering-generative-interface |
Local HTML interface |
| Feedback collection | clawrec-collecting-display-feedback |
Feedback packets |
The detailed artifact boundaries are documented in
docs/architecture.md.
.
├── skills/ # ClawRec method implementation
├── clawrec-simbench/
│ ├── review_md/ # Simulated intent specifications
│ └── cases/ # GUI-grounded event cases
├── evaluation/
│ ├── prompts/ # Versioned LLM-judge prompts
│ ├── scs/ # User-state evaluation
│ └── trajectory/ # Final recommendation evaluation
└── docs/
- Git
- Python 3.9 or newer for evaluation utilities
- An OpenClaw-compatible agent runtime with Skill support
- A visible, user-authorized browser session for live observation and search
The evaluation scripts use only the Python standard library.
git clone <REPOSITORY_URL> ClawRec
cd ClawRecWith an OpenClaw CLI that exposes Skill installation:
for skill_dir in skills/clawrec-*; do
openclaw skills install "$skill_dir"
done
openclaw skills checkIf the runtime uses a different installation command, register every directory
under skills/ without changing its internal file structure.
Run commands from the workspace where ClawRec should create its .clawrec/
state.
- Open a visible browser session and sign in only to sources you authorize ClawRec to use.
- Start the OpenClaw interactive interface:
openclaw tui- Enter the following prompt in the TUI:
Use clawrec-using-clawrec to run one full ClawRec recommendation cycle.
The lifecycle routes through observation, state synthesis, retrieval planning, search execution, curation, and interface rendering. Headless browsing is not used.
Start openclaw tui, then enter:
Use clawrec-using-clawrec to refresh my ClawRec user state from browser behavior I authorize.
Start openclaw tui, then enter:
Use clawrec-using-clawrec to generate a recommendation page from the current ClawRec user state.
A successful full cycle writes the following artifacts under the active workspace:
.clawrec/
├── clawrec-unified-events/
│ └── latest/
│ ├── evidence.jsonl
│ └── meta.json
├── clawrec-user-state/
│ ├── USER_STATE.md
│ ├── STATE_UPDATES.md
│ ├── NEGATIVE_SIGNALS.md
│ └── latest-update-packets.md
├── clawrec-content-plan/
│ └── latest/
│ └── PLAN.md
├── clawrec-search-execution/
│ └── latest/
│ ├── RUN.md
│ ├── tasks.jsonl
│ ├── results.jsonl
│ ├── errors.jsonl
│ └── meta.json
├── clawrec-display-data/
│ └── latest/
│ ├── display-data.json
│ ├── candidates.jsonl
│ ├── excluded.jsonl
│ ├── CURATION.md
│ └── meta.json
└── clawrec-generative-interface/
└── <session_id>/
├── content/
│ └── *.html
└── state/
├── displayed-cards.json
└── events/
The main rendering input, display-data.json, has sectioned recommendation
cards:
{
"sections": [
{
"label": "Section label",
"topic": "Recommendation topic",
"cards": [
{
"title": "Item title",
"source": "Source name",
"href": "https://example.com/item",
"summary": "Why this item is useful now"
}
]
}
]
}After explicit interaction collection, feedback artifacts appear under:
.clawrec/clawrec-feedback/latest/
├── feedback.jsonl
├── section-feedback.jsonl
├── feedback-packets.md
├── COLLECTION.md
└── meta.json
clawrec-simbench/ contains the simulated intent
specifications and GUI-grounded event cases used by the evaluation protocol.
Each case separates:
input/: evidence and constraints visible to the system;private/: held-out targets available only to the evaluator.
To create writable runtime cases without exposing evaluator-only files:
mkdir -p .local/cases
rsync -a --exclude='private/' \
clawrec-simbench/cases/ \
.local/cases/Process events in numeric order within each workspace and preserve that
workspace's .clawrec/ state between events. See
docs/reproduction.md for the experiment boundary.
The repository provides two OpenAI-compatible LLM-judge pipelines.
python3 evaluation/scs/build_scs_judge_inputs.py \
--cases-dir .local/cases
OPENAI_API_KEY=... python3 evaluation/scs/run_scs_llm_judge.py \
--model deepseek-v4-flash \
--concurrency 4
python3 evaluation/scs/aggregate_scs_judgments.pyExpected aggregate outputs:
.local/artifacts/eval/scs/runs/<run_id>/
├── judge_inputs.jsonl
├── input_problems.jsonl
├── input_manifest.json
├── judgments.jsonl
├── raw_responses.jsonl
├── judge_errors.jsonl # Present when requests or responses fail
├── scores.csv
└── summary.json
python3 evaluation/trajectory/build_trajectory_judge_inputs.py \
--cases-dir .local/cases \
--k 20
OPENAI_API_KEY=... python3 evaluation/trajectory/run_trajectory_llm_judge.py \
--model deepseek-v4-flash \
--concurrency 4
python3 evaluation/trajectory/aggregate_trajectory_judgments.pyExpected aggregate outputs:
.local/artifacts/eval/trajectory/runs/<run_id>/
├── judge_inputs.jsonl
├── input_problems.jsonl
├── input_manifest.json
├── judgments.jsonl
├── raw_responses.jsonl
├── judge_errors.jsonl # Present when requests or responses fail
├── card_scores.csv
├── scores.csv
└── summary.json
Set OPENAI_BASE_URL, OPENAI_API_KEY, and OPENAI_MODEL to use another
OpenAI-compatible endpoint. Detailed options are listed in
evaluation/README.md.
ClawRec achieves an NDCG@20 of 0.6134 and a Hit@20 of 0.6944 on ClawRec-SimBench. See the paper for comparisons, ablations, and diagnostics.