Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

lex-chile-embeddings

Article-level embeddings of the complete Chilean legal corpus.

Pre-computed vectors for semantic search, RAG, clustering, and legal NLP research. Auto-synced weekly with the source markdown repository.

Dataset

Field Description
id Unique row identifier
law_name Law name / number (e.g., "Código Civil")
article_number Article identifier
chunk_idx Chunk index within article (0 if article fits in one chunk)
text Raw article text
embedding 1024-dim float vector
text_hash SHA-256 of text (used for change detection)
status active or repealed
token_count Token count (model tokenizer)
last_modified Timestamp of last change
source_url Link to source markdown file

Model: intfloat/multilingual-e5-large (1024 dimensions) Format: Parquet Published at: huggingface.co/datasets/on1link/lex-chile-embeddings [TODO]

Usage

from datasets import load_dataset
import numpy as np

ds = load_dataset("on1link/lex-chile-embeddings", split="train")

# Load into FAISS
embeddings = np.array(ds["embedding"], dtype="float32")

import faiss
index = faiss.IndexFlatIP(embeddings.shape[1])
faiss.normalize_L2(embeddings)
index.add(embeddings)

# Query
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/multilingual-e5-large")
q = model.encode(["query: derecho a vacaciones laborales"], normalize_embeddings=True)
scores, indices = index.search(q, k=5)

for i in indices[0]:
    print(f"{ds[int(i)]['law_name']} Art. {ds[int(i)]['article_number']}")
    print(ds[int(i)]["text"][:200], "\n")

Auto-sync pipeline

A GitHub Actions workflow runs weekly:

  1. Checks source repo for new commits
  2. Parses markdown → structured records
  3. Computes SHA-256 per article, diffs against current dataset
  4. Re-embeds only changed/new articles
  5. Marks removed articles as status: repealed
  6. Pushes updated parquet to HuggingFace Hub
  7. Updates CHANGELOG.md
┌─────────────────┐     ┌──────────────┐     ┌──────────────┐
│ Source repo (MD) │────►│ Parse + Diff │────►│ Embed deltas │
└─────────────────┘     └──────────────┘     └──────┬───────┘
                                                     │
                                                     ▼
                                              ┌──────────────┐
                                              │  Push to HF   │
                                              └──────────────┘

Freshness guarantee: synced weekly. Last sync commit SHA and timestamp are recorded in this README automatically by the pipeline.

Project structure

lex-chile-embeddings/
├── pipeline/
│   ├── parse.py           # Markdown → structured records
│   ├── diff.py            # Hash-based change detection
│   ├── embed.py           # Batch embedding with sentence-transformers
│   └── publish.py         # Push to HuggingFace Hub
├── .github/
│   └── workflows/
│       └── sync.yml       # Weekly cron workflow
├── CHANGELOG.md
├── LICENSE
└── README.md

Related projects

Repo Description
lex-chile RAG system powered by this dataset
lex-bench QA evaluation benchmark

License

Code: MIT — see LICENSE. Dataset: CC-BY-4.0.

The underlying legal texts are public domain under Chilean law (Art. 7, Ley 17.336). The markdown source is subject to the license of its originating repository. This dataset adds a transformative layer (embeddings, structured metadata, sync pipeline) and is distributed under CC-BY-4.0.

About

Pre-computed article-level embeddings of Chilean law · auto-synced weekly · HuggingFace dataset

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors