Article-level embeddings of the complete Chilean legal corpus.
Pre-computed vectors for semantic search, RAG, clustering, and legal NLP research. Auto-synced weekly with the source markdown repository.
| Field | Description |
|---|---|
id |
Unique row identifier |
law_name |
Law name / number (e.g., "Código Civil") |
article_number |
Article identifier |
chunk_idx |
Chunk index within article (0 if article fits in one chunk) |
text |
Raw article text |
embedding |
1024-dim float vector |
text_hash |
SHA-256 of text (used for change detection) |
status |
active or repealed |
token_count |
Token count (model tokenizer) |
last_modified |
Timestamp of last change |
source_url |
Link to source markdown file |
Model: intfloat/multilingual-e5-large (1024 dimensions)
Format: Parquet
Published at: huggingface.co/datasets/on1link/lex-chile-embeddings [TODO]
from datasets import load_dataset
import numpy as np
ds = load_dataset("on1link/lex-chile-embeddings", split="train")
# Load into FAISS
embeddings = np.array(ds["embedding"], dtype="float32")
import faiss
index = faiss.IndexFlatIP(embeddings.shape[1])
faiss.normalize_L2(embeddings)
index.add(embeddings)
# Query
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/multilingual-e5-large")
q = model.encode(["query: derecho a vacaciones laborales"], normalize_embeddings=True)
scores, indices = index.search(q, k=5)
for i in indices[0]:
print(f"{ds[int(i)]['law_name']} Art. {ds[int(i)]['article_number']}")
print(ds[int(i)]["text"][:200], "\n")A GitHub Actions workflow runs weekly:
- Checks source repo for new commits
- Parses markdown → structured records
- Computes SHA-256 per article, diffs against current dataset
- Re-embeds only changed/new articles
- Marks removed articles as
status: repealed - Pushes updated parquet to HuggingFace Hub
- Updates
CHANGELOG.md
┌─────────────────┐ ┌──────────────┐ ┌──────────────┐
│ Source repo (MD) │────►│ Parse + Diff │────►│ Embed deltas │
└─────────────────┘ └──────────────┘ └──────┬───────┘
│
▼
┌──────────────┐
│ Push to HF │
└──────────────┘
Freshness guarantee: synced weekly. Last sync commit SHA and timestamp are recorded in this README automatically by the pipeline.
lex-chile-embeddings/
├── pipeline/
│ ├── parse.py # Markdown → structured records
│ ├── diff.py # Hash-based change detection
│ ├── embed.py # Batch embedding with sentence-transformers
│ └── publish.py # Push to HuggingFace Hub
├── .github/
│ └── workflows/
│ └── sync.yml # Weekly cron workflow
├── CHANGELOG.md
├── LICENSE
└── README.md
| Repo | Description |
|---|---|
lex-chile |
RAG system powered by this dataset |
lex-bench |
QA evaluation benchmark |
Code: MIT — see LICENSE. Dataset: CC-BY-4.0.
The underlying legal texts are public domain under Chilean law (Art. 7, Ley 17.336). The markdown source is subject to the license of its originating repository. This dataset adds a transformative layer (embeddings, structured metadata, sync pipeline) and is distributed under CC-BY-4.0.