Prerequisites
Install Ollama
Ollama must be running locally before using this provider.1
Download and Install
Download Ollama from ollama.ai and follow the installation instructions for your operating system:
- macOS: Download and run the installer
- Linux: Run
curl -fsSL https://ollama.ai/install.sh | sh - Windows: Download the Windows installer
2
Start Ollama Service
After installation, start the Ollama service:By default, Ollama runs on
http://localhost:114343
Pull a Model
Basic Usage
Configuration Options
Required Options
string
required
Ollama model name to use for evaluations.Examples:
llama3.2, mistral, codellama, gemmaThe model must already be pulled via
ollama pull <model-name>Optional Options
string
default:"http://localhost:11434"
Ollama server base URL. Change this if Ollama is running on a different host or port.Environment variable:
OLLAMA_BASE_URLExample: --base-url http://192.168.1.100:11434boolean
Return log probabilities for each token in the response.
Model Options
Ollama supports extensive model configuration through the following parameters:float
default:"0.8"
Model temperature - higher values make answers more creative.Range: 0.0 to 2.0
integer
default:"40"
Reduces probability of generating nonsense. Higher values give more diverse answers.
float
default:"0.9"
Works with top-k. Higher values lead to more diverse text.Range: 0.0 to 1.0
integer
default:"128"
Maximum number of tokens to predict.Special values:
-1: Infinite generation-2: Fill context window
integer
default:"2048"
Size of the context window (number of tokens).
float
default:"1.1"
How strongly to penalize repetitions. Higher values reduce repetition.
integer
default:"64"
How far back to look to prevent repetition.Special values:
0: Disabled-1: Usenum_ctxvalue
integer
default:"0"
Random number seed for generation. Use the same seed for reproducible outputs.
string[]
Stop sequences - generation stops when these strings are encountered.Example:
--stop END --stop STOPfloat
default:"1"
Tail free sampling - reduces impact of less probable tokens.
integer
default:"0"
Enable Mirostat sampling for controlling perplexity.Options:
0: Disabled1: Mirostat 1.02: Mirostat 2.0
float
default:"5.0"
Mirostat tau - controls balance between coherence and diversity.
float
default:"0.1"
Mirostat learning rate.
Hardware Options
integer
Number of layers to send to GPU(s). Use to control GPU memory usage.
integer
Number of threads to use during computation. Adjust based on your CPU cores.
integer
Number of GQA (Grouped Query Attention) groups in transformer layer. Model-specific setting.
Examples
Basic Single-Turn Evaluation
Multi-Turn with Custom Temperature
Remote Ollama Instance
Reproducible Results with Seed
Large Context Window Configuration
GPU Optimization
Advanced Sampling Configuration
Popular Models
Here are some popular models available through Ollama:For a complete list of available models, visit the Ollama Library.
Environment Variables
Tips
Troubleshooting
Connection Issues
If you see connection errors:- Verify Ollama is running:
ollama list - Check the service is accessible:
curl http://localhost:11434 - Ensure the model is pulled:
ollama pull <model-name>
Performance Issues
- Use
--num-threadto match your CPU cores - Adjust
--num-gputo optimize GPU usage - Consider using smaller models for faster evaluations