English | 日本語
ollama_demo_zoom_en.mp4
![]() |
![]() |
![]() |
A benchmark tool for comparing performance of multiple Ollama models in Google Colab environment. Utilizes free T4 GPU to measure generation speed, response time, memory usage, and other metrics.
- Checkbox-based model selection UI
- Comprehensive performance metrics (speed, TTFT, model size, etc.)
- Auto-save to Google Drive (JSON format)
- Visualization with graphs and tables
- Model size caching
- Developers who need quantitative model comparisons
- Engineers selecting optimal models for projects
- Researchers conducting benchmarks without paid APIs
English version:
https://colab.research.google.com/github/hiroaki-com/ollama-llm-benchmark/blob/main/ollama_multi_model_benchmarker_en.ipynb
- Open notebook in Google Colab
- Runtime > Change runtime type > Select T4 GPU
- Run Model Registry cell to load model list
- Select test targets via checkboxes
- Run Benchmarker cell
- Save results to Google Drive (optional)
Specify test target models in comma-separated format in Model Registry cell.
model_list = "qwen3:8b, qwen3:14b, qwen2.5-coder:7b, ministral-3:8b"Recommended Model Size for T4 GPU
| Size | Performance | Note |
|---|---|---|
| 8B | Fast | Recommended |
| 14B | Medium | Usable |
| 20B+ | Slow | Not recommended |
Verify model names at https://ollama.com/search.
Configure parameters in Benchmarker cell.
save_to_drive = True # Save to Google Drive
timeout_seconds = 1000 # Timeout (seconds)
custom_test_prompt = "" # Custom prompt (default if empty)Custom Prompt Examples
# Coding task
custom_test_prompt = "Write a Python function to calculate Fibonacci sequence"
# Summarization
custom_test_prompt = "Summarize the following article in 3 sentences..."Results are displayed after benchmark execution.
Top Performance by Category
| Category | Model | Score |
|---|---|---|
| Fastest Generation | qwen3:8b | 45.23 t/s |
| Most Responsive | ministral-3:8b | 0.12 s |
| Quickest Pull | qwen2.5-coder:7b | 23.4 s |
Detailed Metrics
| Model | Speed | TTFT | Total | Tok | Pull | Load | Size |
|---|---|---|---|---|---|---|---|
| qwen3:8b | 45.23 t/s | 0.15s | 12.3s | 500 | 45.2s | 2.1s | 4.7GB |
Saved Files
Google Drive/MyDrive/ollama_benchmark/
├── benchmark_results.json # Integrated results
├── archives/
│ └── YYYYMMDD_HHMMSS_session.json # Session-specific
└── model_size_cache.json # Model size cache
| Metric | Description | Unit |
|---|---|---|
| Generation Speed | Token generation rate | tokens/sec |
| TTFT | Time To First Token | seconds |
| Total Time | Total processing time | seconds |
| Tokens | Number of generated tokens | - |
| Pull Time | Model download time | seconds |
| Load Time | VRAM loading time | seconds |
| Size | Model disk/VRAM size | GB |
Data collected and saved during benchmark execution:
Performance Metrics
tokens_per_sec: Token generation speed (tokens/second)first_token_time: Time to first token (seconds)total_time: Total time from prompt submission to completion (seconds)tokens: Total number of generated tokens
Model Information
model_name: Official model name (e.g., "qwen3:8b")model_size_gb: Model disk size (in GB)quantization: Quantization level (extracted from model name)parameter_size: Parameter count (e.g., "8B", "14B")
System Metrics
pull_time: Model download time (seconds)model_load_time: VRAM loading time (seconds)response: Actual response text from modeltimestamp: Measurement execution timestamp
Session Information
session_id: Session identifier (YYYYMMDD_HHMMSS format)test_prompt: Prompt used for testinggpu_type: GPU used (typically "Tesla T4")python_version: Python version
Write a recursive Python function with type hints and a docstring to compute
the factorial of a number, test it with n = 5, and show only the code and the
expected result.
Customizable via custom_test_prompt parameter.
- Runtime: Google Colab (Python 3.10+)
- LLM Engine: Ollama
- Visualization: matplotlib, IPython
- UI: ipywidgets
- Storage: Google Drive API
Measured model sizes are cached to reduce execution time on subsequent runs.
{
"qwen3:8b": 4.66,
"ministral-3:8b": 5.12
}timeout_seconds = 2000 # For larger models or complex promptsAutomatically checks available disk space before execution.
- Safety margin: 2GB
- Minimum free space for unknown models: 20GB
MIT License. See LICENSE for details.
- Ollama - Local LLM execution engine
- Google Colab - Free GPU environment
- Bug reports: Issues
- Questions & discussions: Discussions


