Skip to content

Commit 8af4158

Browse files
Add compute_sites.csv (the megawatts sheet) + Actions-based Pages deploy
- data/compute_sites.csv: research-tier-only extract with structured compute columns (facility_type, gpu_count, gpu_type, compute_note, power_mw, energy_note); under_construction records grouped as the in-progress category. Generated by pipeline/build_compute_sheet.py from sites_seed.csv + curated verified specs; 4 sync/enum tests. - .github/workflows/pages.yml: deploy docs/ via actions/deploy-pages; Pages build_type switched to workflow — the legacy branch builder hung on 3 of 4 builds today. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1 parent 6607f1b commit 8af4158

4 files changed

Lines changed: 296 additions & 0 deletions

File tree

.github/workflows/pages.yml

Lines changed: 30 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,30 @@
1+
name: Deploy map to GitHub Pages
2+
3+
on:
4+
push:
5+
branches: [main]
6+
workflow_dispatch:
7+
8+
permissions:
9+
contents: read
10+
pages: write
11+
id-token: write
12+
13+
concurrency:
14+
group: pages
15+
cancel-in-progress: true
16+
17+
jobs:
18+
deploy:
19+
runs-on: ubuntu-latest
20+
environment:
21+
name: github-pages
22+
url: ${{ steps.deployment.outputs.page_url }}
23+
steps:
24+
- uses: actions/checkout@v4
25+
- uses: actions/configure-pages@v5
26+
- uses: actions/upload-pages-artifact@v3
27+
with:
28+
path: docs
29+
- id: deployment
30+
uses: actions/deploy-pages@v4

data/compute_sites.csv

Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,18 @@
1+
site_id,firm,site_name,facility_type,city,country,lat,lon,coord_precision,status,gpu_count,gpu_type,compute_note,power_mw,energy_note,confidence,evidence_url
2+
T3-003,Citadel Securities,Citadel Securities on Google Cloud,cloud,Miami FL,USA,25.76600,-80.19243,symbolic,active,,,Research platform on Google Cloud: 'more than 1 million cores concurrently' (joint case study). No owned facility in the public record.,,,confirmed,https://cloud.google.com/transform/citadel-securities-reimagine-quantitative-reseach-cloud-scale-speed
3+
T3-014,G-Research,G-Research ML compute (cluster locations undisclosed; London research lab),undisclosed,London,United Kingdom,51.5162,-0.1301,symbolic,active,,,Own job ads describe 'a global estate of data centres'; locations and scale undisclosed.,,,inferred,http://web.archive.org/web/20250314223052/https://www.gresearch.com/vacancy/R2823-Data-Centre-Engineer/
4+
T3-005,High-Flyer Quant (幻方量化),"High-Flyer Fire-Flyer 2 (萤火二号) AI cluster — undisclosed location, symbolic pin at Hangzhou HQ",self_build,Hangzhou,China,30.2741,120.1551,symbolic,active,10000,NVIDIA A100 (PCIe),"Fire-Flyer 2 (2021): ~10,000 PCIe A100s, ~1,250 GPU nodes, ~200 storage servers, RMB 1bn; trained DeepSeek's early models. Predecessor Fire-Flyer 1 (2020, 1,100 A100s) historical.",3.0,Operator's SC24 paper: total draw 'approximately just over 3 MW' (<4 MW).,confirmed,https://arxiv.org/pdf/2408.14158
5+
T3-011,Hudson River Trading,Hudson River Trading AI research data center — Lefdal Mine Datacenter (Norway),colo,"Lefdal, Stad municipality (near Måløy)",Norway,61.932,5.5062,approximate,active,,NVIDIA HGX B200 (Dell PowerEdge XE9685L),AI research deployment at Lefdal Mine Datacenter: direct-liquid-cooled racks >125kW per cabinet (HRT's own posts).,,Lefdal campus runs on Norwegian hydro; fjord-water cooling.,confirmed,https://www.linkedin.com/posts/hudson-river-trading_hpc-at-hrt-scaling-through-the-fjord-activity-7438251239609516032-ZIEs
6+
T3-012,Hudson River Trading,"Hudson River Trading cloud research compute — Google Cloud + Lambda (symbolic pin, New York)",cloud,New York,USA,40.7128,-74.006,symbolic,active,,NVIDIA GPUs (Google Cloud); HGX B200 (Lambda),Google Cloud HPC for research/simulation (2024); Lambda AI-cloud partnership (2026).,,,confirmed,https://www.googlecloudpresscorner.com/2024-07-18-Google-Cloud-Enables-Hudson-River-Tradings-Automated-Trading-Evolution-with-High-Performance-Compute-and-AI-Infrastructure
7+
T3-009,Jane Street,"Jane Street AI training data center (Dallas, TX — city-level)",undisclosed,Dallas,United States,32.7767,-96.797,city_level,active,4032,liquid-cooled GPUs (type undisclosed),"4,032 liquid-cooled GPUs per Jane Street's careers page (public video tour, May 2026); firmwide 'tens of thousands' of high-end GPUs.",,,reported,https://www.datacenterdynamics.com/en/news/quant-trading-firm-jane-street-plans-data-center-report/
8+
T3-010,Jane Street,"Jane Street on CoreWeave AI cloud (symbolic pin, NYC HQ)",cloud,New York,United States,40.7147,-74.0154,symbolic,active,,NVIDIA Vera Rubin (via CoreWeave),"$6bn AI-cloud agreement with CoreWeave (Apr 2026), next-gen compute 'across multiple facilities', plus $1bn equity investment.",,,confirmed,https://www.coreweave.com/news/jane-street-signs-6-billion-ai-cloud-agreement-with-coreweave
9+
T3-015,Jump Trading,"Jump Trading AI research cluster (Vera Rubin NVL72 — symbolic pin, Chicago HQ)",undisclosed,Chicago,USA,41.8965,-87.644,symbolic,active,,NVIDIA Vera Rubin NVL72,Early Vera Rubin NVL72 rack-scale deployment (~Apr 2026); infrastructure supports 'thousands of CPUs and GPUs' (VAST case study). Cluster location undisclosed.,,,confirmed,https://resources.nvidia.com/en-us-financial-services-industry/jump-trading
10+
T3-016,Jump Trading,"Jump Trading HPC data center (Carrollton, TX — city-level, from job listings)",undisclosed,Carrollton,USA,32.9756,-96.89,city_level,active,,,"HPC data centre in Carrollton, TX signalled by on-site technician job ads (high-density liquid-cooled cabinets); operator and capacity undisclosed.",,,inferred,https://boards-api.greenhouse.io/v1/boards/jumptrading/jobs/7781600
11+
T3-008,Minghong Investment (明汯投资),Minghong Investment HPC cluster (明汯投资) — symbolic pin at Shanghai HQ,undisclosed,Shanghai (HQ city; cluster location undisclosed),China,31.2304,121.4737,symbolic,active,,,"Firm statement (Feb 2025): thousands of GPU cards, tens of thousands of CPU cores, multi-PB storage, ~400 PFlops claimed.",,,reported,https://www.cls.cn/detail/1953863
12+
T3-004,Renaissance Technologies,Renaissance Technologies East Setauket campus,self_build,East Setauket NY,USA,40.93068,-73.11060,exact,active,,,On-prem compute at East Setauket campus long-reported (Zuckerman); scale entirely undisclosed — the absence of any public figure is the finding.,,,inferred,https://reports.adviserinfo.sec.gov/reports/ADV/106661/PDF/106661.pdf
13+
T3-013,Two Sigma,"Two Sigma private datacenters + Google Cloud (symbolic pin, New York HQ)",hybrid,New York,United States,40.7247,-74.0048,symbolic,active,,,"Two private datacenters: 1,313 nodes / 31,512 CPU cores / 328 TB RAM as of 2016 (CMU ATLAS traces, firm co-authored); Google Cloud since (engineering blog).",,,confirmed,https://www.pdl.cmu.edu/PDL-FTP/CloudComputing/cloudstudy_atc18.pdf
14+
T3-007,Ubiquant (九坤投资),Ubiquant 'Bei Ming' AI supercomputing cluster (北溟超算集群) — symbolic pin at Beijing HQ,undisclosed,Beijing (HQ city; cluster location undisclosed),China,39.9042,116.4074,symbolic,active,,,'Bei Ming' AI supercomputing cluster; >RMB 100m invested in 2020 per GM interview. Location and scale undisclosed.,,,reported,https://app-web.chnfund.com/yhrw/gdft/202104/t20210425_3115051.html
15+
T3-001,XTX Markets,XTX Markets Kajaani campus,self_build,Kajaani,Finland,64.22312,27.68174,approximate,active,,,"Campus building 1: 15,000 sqm, 22.5MW IT. Firmwide fleet 25k+ GPUs (Jan 2025 announcement). ~250MW full-build figure is DCD-reported, not in XTX's own materials.",22.5,,confirmed,https://files.xtxmarkets.com/publications/kajaani/index.html
16+
T3-002,XTX Markets (tenant),Verne Global Keflavik campus (XTX Markets + HPC tenants),colo,Keflavik,Iceland,63.9778,-22.5758,exact,active,,,"XTX's pre-Kajaani research cluster at Verne's multi-tenant campus; XTX-specific footprint undisclosed. power_mw is the CAMPUS capacity (>140MW), not XTX's share.",140,Campus runs on 100% hydro/geothermal per operator.,confirmed,https://www.verne.co/iceland
17+
T3-006,DeepSeek (High-Flyer lineage),"DeepSeek Ulanqab data centre (Inner Mongolia, signalled by job postings)",self_build,Ulanqab,China,41.02,113.11,city_level,under_construction,,,"Self-build signalled by official job ads for data-centre delivery/O&M roles in Ulanqab (Apr 2026); no size, MW, or address public.",,"IT之家 cites Ulanqab's cheap power, 4.3°C average temperature, PUE below 1.26 as siting rationale.",inferred,https://www.scmp.com/tech/big-tech/article/3349741/deepseek-ramps-hiring-ahead-v4-launch-questions-swirl-over-chip-strategy
18+
T3-017,XTX Markets,"XTX Markets Kajaani — Building 2 (second data center, Sokajärventie)",self_build,Kajaani,Finland,64.22312,27.68174,approximate,under_construction,,,Campus building 2; IT load and floor area not yet public (building 1 = 22.5MW IT for scale). Up to five buildings originally planned.,,,confirmed,https://www.yitgroup.com/en/news-repository/press-release/yit-continues-collaboration-with-xtx-markets-in-kajaani

pipeline/build_compute_sheet.py

Lines changed: 212 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,212 @@
1+
"""
2+
Build data/compute_sites.csv — the megawatts sheet.
3+
4+
The focused deliverable behind the essay's thesis: ONLY the data/compute
5+
centres of quant/HFT firms (tier=research), with structured compute columns
6+
(gpu_count, gpu_type), power_mw, and an energy_note, split by status
7+
(active vs under_construction).
8+
9+
Source of truth: data/sites_seed.csv supplies identity, coordinates, status,
10+
power_mw, confidence and evidence_url for every research-tier record.
11+
COMPUTE_SPECS below adds the compute-specific fields, hand-curated from the
12+
verifier-checked evidence in data/raw/enrichment_results.json and
13+
data/raw/compute_specs_facts.json (every figure traces to a fetched URL).
14+
15+
No network access. Usage:
16+
python pipeline/build_compute_sheet.py
17+
"""
18+
19+
from __future__ import annotations
20+
21+
from pathlib import Path
22+
23+
import pandas as pd
24+
25+
ROOT = Path(__file__).resolve().parent.parent
26+
SITES_CSV = ROOT / "data" / "sites_seed.csv"
27+
OUT_CSV = ROOT / "data" / "compute_sites.csv"
28+
29+
# facility_type: self_build / colo / cloud / hybrid / undisclosed
30+
COMPUTE_SPECS = {
31+
"T3-001": {
32+
"firm": "XTX Markets",
33+
"facility_type": "self_build",
34+
"gpu_count": "",
35+
"gpu_type": "",
36+
"compute_note": "Campus building 1: 15,000 sqm, 22.5MW IT. Firmwide fleet 25k+ GPUs (Jan 2025 announcement). ~250MW full-build figure is DCD-reported, not in XTX's own materials.",
37+
"energy_note": "",
38+
},
39+
"T3-017": {
40+
"firm": "XTX Markets",
41+
"facility_type": "self_build",
42+
"gpu_count": "",
43+
"gpu_type": "",
44+
"compute_note": "Campus building 2; IT load and floor area not yet public (building 1 = 22.5MW IT for scale). Up to five buildings originally planned.",
45+
"energy_note": "",
46+
},
47+
"T3-002": {
48+
"firm": "XTX Markets (tenant)",
49+
"facility_type": "colo",
50+
"gpu_count": "",
51+
"gpu_type": "",
52+
"compute_note": "XTX's pre-Kajaani research cluster at Verne's multi-tenant campus; XTX-specific footprint undisclosed. power_mw is the CAMPUS capacity (>140MW), not XTX's share.",
53+
"energy_note": "Campus runs on 100% hydro/geothermal per operator.",
54+
},
55+
"T3-003": {
56+
"firm": "Citadel Securities",
57+
"facility_type": "cloud",
58+
"gpu_count": "",
59+
"gpu_type": "",
60+
"compute_note": "Research platform on Google Cloud: 'more than 1 million cores concurrently' (joint case study). No owned facility in the public record.",
61+
"energy_note": "",
62+
},
63+
"T3-004": {
64+
"firm": "Renaissance Technologies",
65+
"facility_type": "self_build",
66+
"gpu_count": "",
67+
"gpu_type": "",
68+
"compute_note": "On-prem compute at East Setauket campus long-reported (Zuckerman); scale entirely undisclosed — the absence of any public figure is the finding.",
69+
"energy_note": "",
70+
},
71+
"T3-005": {
72+
"firm": "High-Flyer Quant (幻方量化)",
73+
"facility_type": "self_build",
74+
"gpu_count": "10000",
75+
"gpu_type": "NVIDIA A100 (PCIe)",
76+
"compute_note": "Fire-Flyer 2 (2021): ~10,000 PCIe A100s, ~1,250 GPU nodes, ~200 storage servers, RMB 1bn; trained DeepSeek's early models. Predecessor Fire-Flyer 1 (2020, 1,100 A100s) historical.",
77+
"energy_note": "Operator's SC24 paper: total draw 'approximately just over 3 MW' (<4 MW).",
78+
},
79+
"T3-006": {
80+
"firm": "DeepSeek (High-Flyer lineage)",
81+
"facility_type": "self_build",
82+
"gpu_count": "",
83+
"gpu_type": "",
84+
"compute_note": "Self-build signalled by official job ads for data-centre delivery/O&M roles in Ulanqab (Apr 2026); no size, MW, or address public.",
85+
"energy_note": "IT之家 cites Ulanqab's cheap power, 4.3°C average temperature, PUE below 1.26 as siting rationale.",
86+
},
87+
"T3-007": {
88+
"firm": "Ubiquant (九坤投资)",
89+
"facility_type": "undisclosed",
90+
"gpu_count": "",
91+
"gpu_type": "",
92+
"compute_note": "'Bei Ming' AI supercomputing cluster; >RMB 100m invested in 2020 per GM interview. Location and scale undisclosed.",
93+
"energy_note": "",
94+
},
95+
"T3-008": {
96+
"firm": "Minghong Investment (明汯投资)",
97+
"facility_type": "undisclosed",
98+
"gpu_count": "",
99+
"gpu_type": "",
100+
"compute_note": "Firm statement (Feb 2025): thousands of GPU cards, tens of thousands of CPU cores, multi-PB storage, ~400 PFlops claimed.",
101+
"energy_note": "",
102+
},
103+
"T3-009": {
104+
"firm": "Jane Street",
105+
"facility_type": "undisclosed",
106+
"gpu_count": "4032",
107+
"gpu_type": "liquid-cooled GPUs (type undisclosed)",
108+
"compute_note": "4,032 liquid-cooled GPUs per Jane Street's careers page (public video tour, May 2026); firmwide 'tens of thousands' of high-end GPUs.",
109+
"energy_note": "",
110+
},
111+
"T3-010": {
112+
"firm": "Jane Street",
113+
"facility_type": "cloud",
114+
"gpu_count": "",
115+
"gpu_type": "NVIDIA Vera Rubin (via CoreWeave)",
116+
"compute_note": "$6bn AI-cloud agreement with CoreWeave (Apr 2026), next-gen compute 'across multiple facilities', plus $1bn equity investment.",
117+
"energy_note": "",
118+
},
119+
"T3-011": {
120+
"firm": "Hudson River Trading",
121+
"facility_type": "colo",
122+
"gpu_count": "",
123+
"gpu_type": "NVIDIA HGX B200 (Dell PowerEdge XE9685L)",
124+
"compute_note": "AI research deployment at Lefdal Mine Datacenter: direct-liquid-cooled racks >125kW per cabinet (HRT's own posts).",
125+
"energy_note": "Lefdal campus runs on Norwegian hydro; fjord-water cooling.",
126+
},
127+
"T3-012": {
128+
"firm": "Hudson River Trading",
129+
"facility_type": "cloud",
130+
"gpu_count": "",
131+
"gpu_type": "NVIDIA GPUs (Google Cloud); HGX B200 (Lambda)",
132+
"compute_note": "Google Cloud HPC for research/simulation (2024); Lambda AI-cloud partnership (2026).",
133+
"energy_note": "",
134+
},
135+
"T3-013": {
136+
"firm": "Two Sigma",
137+
"facility_type": "hybrid",
138+
"gpu_count": "",
139+
"gpu_type": "",
140+
"compute_note": "Two private datacenters: 1,313 nodes / 31,512 CPU cores / 328 TB RAM as of 2016 (CMU ATLAS traces, firm co-authored); Google Cloud since (engineering blog).",
141+
"energy_note": "",
142+
},
143+
"T3-014": {
144+
"firm": "G-Research",
145+
"facility_type": "undisclosed",
146+
"gpu_count": "",
147+
"gpu_type": "",
148+
"compute_note": "Own job ads describe 'a global estate of data centres'; locations and scale undisclosed.",
149+
"energy_note": "",
150+
},
151+
"T3-015": {
152+
"firm": "Jump Trading",
153+
"facility_type": "undisclosed",
154+
"gpu_count": "",
155+
"gpu_type": "NVIDIA Vera Rubin NVL72",
156+
"compute_note": "Early Vera Rubin NVL72 rack-scale deployment (~Apr 2026); infrastructure supports 'thousands of CPUs and GPUs' (VAST case study). Cluster location undisclosed.",
157+
"energy_note": "",
158+
},
159+
"T3-016": {
160+
"firm": "Jump Trading",
161+
"facility_type": "undisclosed",
162+
"gpu_count": "",
163+
"gpu_type": "",
164+
"compute_note": "HPC data centre in Carrollton, TX signalled by on-site technician job ads (high-density liquid-cooled cabinets); operator and capacity undisclosed.",
165+
"energy_note": "",
166+
},
167+
}
168+
169+
COLUMNS = [
170+
"site_id", "firm", "site_name", "facility_type", "city", "country",
171+
"lat", "lon", "coord_precision", "status", "gpu_count", "gpu_type",
172+
"compute_note", "power_mw", "energy_note", "confidence", "evidence_url",
173+
]
174+
175+
176+
def main() -> None:
177+
sites = pd.read_csv(SITES_CSV, dtype=str).fillna("")
178+
research = sites[sites.tier == "research"].set_index("site_id")
179+
180+
missing = set(research.index) - set(COMPUTE_SPECS)
181+
extra = set(COMPUTE_SPECS) - set(research.index)
182+
if missing or extra:
183+
raise SystemExit(f"COMPUTE_SPECS out of sync: missing={sorted(missing)} extra={sorted(extra)}")
184+
185+
rows = []
186+
for site_id, spec in COMPUTE_SPECS.items():
187+
s = research.loc[site_id]
188+
rows.append({
189+
"site_id": site_id,
190+
"site_name": s.site_name,
191+
"city": s.city, "country": s.country,
192+
"lat": s.lat, "lon": s.lon,
193+
"coord_precision": s.coord_precision,
194+
"status": s.status,
195+
"power_mw": s.power_mw,
196+
"confidence": s.confidence,
197+
"evidence_url": s.evidence_url,
198+
**spec,
199+
})
200+
201+
df = pd.DataFrame(rows)[COLUMNS]
202+
# Under-construction sites grouped last — the "in progress" category.
203+
df["_k"] = df.status.map({"active": 0, "historical": 1, "under_construction": 2})
204+
df = df.sort_values(["_k", "firm", "site_id"]).drop(columns="_k")
205+
df.to_csv(OUT_CSV, index=False)
206+
print(f"Wrote {OUT_CSV.name}: {len(df)} compute sites "
207+
f"({df.status.value_counts().to_dict()}); "
208+
f"power stated for {(df.power_mw != '').sum()} sites")
209+
210+
211+
if __name__ == "__main__":
212+
main()

tests/test_data.py

Lines changed: 36 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -23,6 +23,12 @@
2323
"path_id", "origin_site_id", "dest_site_id", "operator", "medium",
2424
"status", "confidence", "evidence_note", "evidence_url",
2525
]
26+
COMPUTE_COLUMNS = [
27+
"site_id", "firm", "site_name", "facility_type", "city", "country",
28+
"lat", "lon", "coord_precision", "status", "gpu_count", "gpu_type",
29+
"compute_note", "power_mw", "energy_note", "confidence", "evidence_url",
30+
]
31+
FACILITY_TYPES = {"self_build", "colo", "cloud", "hybrid", "undisclosed"}
2632

2733
TIERS = {"execution", "network", "research"}
2834
CONFIDENCE = {"confirmed", "reported", "inferred"}
@@ -114,6 +120,36 @@ def test_path_endpoints_exist(paths, sites):
114120
assert refs <= known, f"paths reference unknown site_ids: {sorted(refs - known)}"
115121

116122

123+
@pytest.fixture(scope="module")
124+
def compute() -> pd.DataFrame:
125+
return pd.read_csv(DATA / "compute_sites.csv", dtype=str).fillna("")
126+
127+
128+
def test_compute_schema(compute):
129+
assert list(compute.columns) == COMPUTE_COLUMNS
130+
131+
132+
def test_compute_in_sync_with_sites(compute, sites):
133+
research = sites[sites.tier == "research"]
134+
assert set(compute.site_id) == set(research.site_id), \
135+
"compute_sites.csv must cover exactly the research-tier records; rerun pipeline/build_compute_sheet.py"
136+
joined = compute.merge(research, on="site_id", suffixes=("_c", ""))
137+
for col in ("status", "confidence", "coord_precision"):
138+
mismatched = joined[joined[f"{col}_c"] != joined[col].astype(str)]
139+
assert mismatched.empty, f"{col} out of sync for: {mismatched.site_id.tolist()}"
140+
141+
142+
def test_compute_enums(compute):
143+
assert compute.facility_type.isin(FACILITY_TYPES).all()
144+
assert compute.status.isin(STATUS).all()
145+
146+
147+
def test_compute_gpu_count_numeric(compute):
148+
nonempty = compute[compute.gpu_count != ""]
149+
assert nonempty.gpu_count.str.fullmatch(r"\d+").all(), \
150+
"gpu_count must be a bare integer or empty"
151+
152+
117153
def test_path_enums_and_urls(paths):
118154
if paths.empty:
119155
return

0 commit comments

Comments
 (0)