Terminal AI agent and coding assistant in the terminal for Ollama, SGLang, xAI, OpenAI, and any OpenAI-compatible API. Full-screen TUI, optional file/shell tools, and Model Context Protocol (MCP) client support—one static Go binary.
Lab 1 in the AI · Agents · MCP learning path · Audience: intermediate · Time: ~1 hour to first chat · Level: intermediate
SEO keywords: terminal AI agent, Ollama TUI, SGLang OpenAI-compatible, OpenAI compatible CLI agent, MCP client Go, local LLM coding agent, agenterm tutorial, AI pair programmer terminal.
| Best for | Developers who want Ollama, SGLang, or any /v1 chat API in the terminal with optional coding tools |
| Runs on | Linux, macOS, Windows · Docker / Podman |
| Not | A web UI or desktop app — pure terminal |
You (TUI)
│
▼
agenterm ──► Ollama / SGLang / xAI / OpenAI (POST /v1/chat/completions)
│
└── tools ──► built-in (files, optional shell) + MCP servers
1. Install (Linux / macOS):
curl -fsSL https://raw.githubusercontent.com/saurabhahuja71/agenterm/main/scripts/install.sh | bashBinary goes to ~/.local/bin/agenterm. If the shell cannot find it:
export PATH="$HOME/.local/bin:$PATH"2. Start a backend (pick one):
# Ollama (default)
ollama pull qwen2.5-coder:32b
ollama serve # http://127.0.0.1:11434
# or SGLang (OpenAI-compatible on :30000 — see SGLang section below)
# sglang serve --model-path /path/to/model.gguf --quantization gguf \
# --host 127.0.0.1 --port 30000 --served-model-name qwen2.5-coder-32b-q4_k_m.gguf3. Chat:
agenterm init # first time: writes ~/.agenterm/config.toml
agenterm --ping # check the API
agenterm # open the TUI (Ollama default)
# SGLang instead:
agenterm --provider sglang --ping
agenterm --provider sglangIn the TUI: type a message and press Enter. Try /help, /model, or /tools off for pure chat.
- Why agenterm?
- Install
- Quick start
- Local vs remote Ollama
- SGLang
- Docker / Podman
- CLI and in-chat commands
- Configuration
- Tools and MCP
- FAQ
- Architecture
- Releases
- License
| You want… | agenterm does… |
|---|---|
| Chat with local LLMs | Defaults to Ollama at http://127.0.0.1:11434/v1 |
| Faster SGLang serving | --provider sglang → http://127.0.0.1:30000/v1 |
| A GPU box over the network | --base-url or an SSH tunnel to localhost |
| A coding agent in the shell | Tool loop: list/read/write files; optional shell; MCP |
| Switch models mid-session | /model while chatting |
| Fast small talk | Greetings skip tools; use --no-tools for pure chat |
| No Python/Node runtime | One static Go binary |
Not affiliated with xAI. “Grok-style” only means a snappy terminal agent UX.
curl -fsSL https://raw.githubusercontent.com/saurabhahuja71/agenterm/main/scripts/install.sh | bash
agenterm --version| Variable | Meaning |
|---|---|
INSTALL_DIR |
Install path (default ~/.local/bin) |
AGENTERM_VERSION |
Release tag or latest |
AGENTERM_FROM_SOURCE=1 |
Build with Go instead of downloading a release |
AGENTERM_SKIP_INIT=1 |
Do not create config |
curl -fsSL -o agenterm \
https://github.com/saurabhahuja71/agenterm/releases/latest/download/agenterm-linux-amd64
chmod +x agenterm && ./agentermAlso published: linux-arm64, darwin-amd64, darwin-arm64, windows-amd64.exe.
git clone https://github.com/saurabhahuja71/agenterm.git
cd agenterm
make build # → ./agenterm
make install # → $(go env GOPATH)/bin (or $GOBIN)
# or:
go install github.com/saurabhahuja71/agenterm/cmd/agenterm@latestollama pull qwen2.5-coder:32b
agenterm --ping
agenterm -m qwen2.5-coder:32b
# pure chat (no function tools):
agenterm --no-tools -m qwen2.5-coder:32bssh -N -L 11434:127.0.0.1:11434 user@gpu-host
# agenterm still uses http://127.0.0.1:11434/v1
agentermSGLang exposes the same OpenAI-compatible /v1/chat/completions surface. Default port in this project’s docs: 30000.
# Server already on this machine (or tunnel already open):
agenterm --provider sglang --ping
agenterm --provider sglang
agenterm --provider sglang --no-tools # pure chatSSH tunnel from a laptop to a GPU host:
ssh -N -L 30000:127.0.0.1:30000 user@gpu-host
# or with a jump host:
# ssh -N -J bastion -L 30000:127.0.0.1:30000 opc@gpu-host
agenterm --provider sglangModel id must match SGLang’s --served-model-name (often the GGUF basename). List with:
curl -s http://127.0.0.1:30000/v1/models
# then:
agenterm --provider sglang -m qwen2.5-coder-32b-q4_k_m.gguf| Action | How |
|---|---|
| Send message | Enter |
| Help | /help |
| List / switch model | /model · /model qwen2.5-coder:32b |
| Disable tools this session | /tools off |
| Clear history | /clear or Ctrl+L |
| Quit | Ctrl+C |
Run agenterm from the project directory you care about (cd myrepo && agenterm) so file tools use that workspace.
provider = "ollama-local"
model = "qwen2.5-coder:32b"
base_url = "http://127.0.0.1:11434/v1"
api_key = "ollama"Remote host on the LAN:
provider = "ollama-remote"
[providers.ollama-remote]
base_url = "http://192.168.1.50:11434/v1"
api_key = "ollama"
model = "qwen2.5"agenterm --base-url http://127.0.0.1:11434/v1 -m qwen2.5-coder:32b
agenterm --base-url http://gpu-box:11434/v1 -m qwen2.5
export XAI_API_KEY=xai-...
agenterm --provider xai -m grok-3
export OPENAI_API_KEY=sk-...
agenterm --provider openai -m gpt-4o-miniAlways include /v1 on Ollama base URLs (/v1/chat/completions, /v1/models).
Full example: configs/config.example.toml.
SGLang is an OpenAI-compatible inference server (often used with GGUF / multi-GPU). agenterm talks to it the same way as Ollama: POST {base_url}/chat/completions.
| Ollama (default) | SGLang | |
|---|---|---|
| Typical URL | http://127.0.0.1:11434/v1 |
http://127.0.0.1:30000/v1 |
| Provider preset | ollama-local / ollama-remote |
sglang |
| Model id | Ollama tag (qwen2.5-coder:32b) |
--served-model-name (e.g. GGUF basename) |
| API key | any string (e.g. ollama) |
any string (e.g. sglang) |
provider = "sglang"
# optional top-level overrides; preset also sets these under [providers.sglang]
[providers.sglang]
base_url = "http://127.0.0.1:30000/v1"
api_key = "sglang"
# Must match the id returned by GET /v1/models
model = "qwen2.5-coder-32b-q4_k_m.gguf"Or keep Ollama as default and only switch on the CLI:
agenterm --provider sglang
agenterm --base-url http://127.0.0.1:30000/v1 -m qwen2.5-coder-32b-q4_k_m.gguf --api-key sglangEnv equivalent:
export AGENTERM_PROVIDER=sglang
export AGENTERM_BASE_URL=http://127.0.0.1:30000/v1
export AGENTERM_MODEL=qwen2.5-coder-32b-q4_k_m.gguf
export AGENTERM_API_KEY=sglang
agenterm --ping && agentermOn the GPU host (adjust paths, tensor-parallel size, and context to your hardware):
# Example: dense Qwen2.5 coder GGUF, OpenAI API on localhost:30000
sglang serve \
--model-path /path/to/qwen2.5-coder-32b-q4_k_m.gguf \
--quantization gguf \
--host 127.0.0.1 \
--port 30000 \
--tp-size 2 \
--context-length 32768 \
--served-model-name qwen2.5-coder-32b-q4_k_m.gguf \
--trust-remote-codePrefer dense GGUF weights for stability. Some MoE GGUFs (e.g. certain qwen3-coder builds) can fail at runtime depending on the SGLang build—use Ollama for those tags if needed.
SGLang is often bound to 127.0.0.1 only on the GPU box (no public ingress). From your workstation:
ssh -N -L 30000:127.0.0.1:30000 user@gpu-host
# jump host:
# ssh -N -J bastion -L 30000:127.0.0.1:30000 opc@gpu-host
curl -s http://127.0.0.1:30000/v1/models
agenterm --provider sglang --ping
agenterm --provider sglangIf you use a corporate HTTP proxy, exclude localhost so the tunnel is not proxied:
export no_proxy="localhost,127.0.0.1,${no_proxy:-}"
export NO_PROXY="localhost,127.0.0.1,${NO_PROXY:-}"Do not load a heavy Ollama model and a heavy SGLang model on the same GPUs at once (memory contention). Unload Ollama keep-alives before starting SGLang, or stop the SGLang service before ollama run of a large tag.
| Goal | Suggested path |
|---|---|
Everyday chat / tags from ollama list |
Ollama · agenterm (default) |
| Lower-latency GGUF serve, multi-GPU TP | SGLang · agenterm --provider sglang |
Always include /v1 on the base URL (/v1/chat/completions, /v1/models).
Prefer the native binary for daily use. Use a container when you want a locked-down runtime.
podman run --rm -it --network=host \
-e AGENTERM_BASE_URL=http://127.0.0.1:11434/v1 \
-v "$HOME/.agenterm:/home/agenterm/.agenterm:Z" \
ghcr.io/saurabhahuja71/agenterm:latestFrom this repo:
./scripts/run-podman.sh --ping
./scripts/run-podman.sh| Need | Why |
|---|---|
-it / TTY |
Interactive TUI |
--network=host (Linux) |
Reach host Ollama (:11434) or SGLang (:30000) |
Volume ~/.agenterm |
Persist config |
On macOS Docker Desktop, use http://host.docker.internal:11434/v1 for host Ollama (or :30000 for host SGLang).
export DOCKER_HOST=unix:///run/user/$(id -u)/podman/podman.sock # rootless Podman
podman compose run --rm agentermagenterm # open TUI
agenterm init # default config (--force to overwrite)
agenterm --ping # connectivity check
agenterm -m qwen2.5 # start with a model
agenterm --provider ollama-remote
agenterm --provider sglang
agenterm --base-url http://host:11434/v1
agenterm --base-url http://127.0.0.1:30000/v1 -m qwen2.5-coder-32b-q4_k_m.gguf
agenterm --no-tools # pure chat (faster)
agenterm --shell # allow run_shell
agenterm --no-mcp/help
/status
/model # list models (* = current)
/model qwen2.5-coder:32b # switch model
/models # alias for /model
/tools on | /tools off
/clear
/quit
Greetings skip tools automatically. For fully tool-free sessions: /tools off, enable_tools = false, or AGENTERM_ENABLE_TOOLS=0.
| Env | Meaning |
|---|---|
AGENTERM_BASE_URL |
API root (e.g. http://127.0.0.1:11434/v1 or http://127.0.0.1:30000/v1) |
AGENTERM_MODEL |
Model id / Ollama tag / SGLang served name |
AGENTERM_PROVIDER |
ollama-local, ollama-remote, sglang, xai, openai, custom |
AGENTERM_ENABLE_TOOLS |
0/false off · 1/true on |
AGENTERM_API_KEY / OLLAMA_API_KEY / XAI_API_KEY / OPENAI_API_KEY |
Auth |
AGENTERM_CONFIG |
Path to config file |
Default file: ~/.agenterm/config.toml.
| Tool | Notes |
|---|---|
list_dir |
List a directory (relative to process cwd) |
read_file |
Read a file |
write_file |
Write a full file |
str_replace |
Surgical edit (preferred for patches) |
find_files |
Locate files when the path is unknown |
git |
Allowlisted git (status, checkout, add, commit, …) |
run_shell |
Off by default; --shell or enable_shell = true |
When you ask to apply / implement / do it, agenterm uses tools—not only printed shell recipes.
agenterm is an MCP client. Example:
[[mcp_servers]]
name = "mcp-demo"
enabled = true
url = "http://127.0.0.1:8080/mcp"Disable for a session: agenterm --no-mcp. Demo server: mcp-demo.
An open-source terminal AI agent in Go: a full-screen TUI that streams chat from Ollama, SGLang, or any OpenAI-compatible API, with optional tools (files, shell, MCP).
No—it is a terminal TUI, not a web or desktop GUI. Use it when you want chat and agent tools without leaving the shell.
Yes. Point base_url / --base-url at the host, or tunnel remote Ollama to 127.0.0.1:11434 and keep the default URL.
Yes. --provider sglang (default http://127.0.0.1:30000/v1), or any custom base URL that exposes /v1/chat/completions. Use the served model name from GET /v1/models, not the Ollama tag, unless you set them to match. See SGLang.
Yes. --provider xai or --provider openai with the matching API key env vars, or any custom --base-url that speaks OpenAI chat completions.
/model to list, then /model <name> (for example /model qwen2.5-coder:32b).
Some models called tools (e.g. list_dir) on greetings. Trivial chat now skips tools; use --no-tools or /tools off for pure chat.
No for release binaries. agenterm is a static Go binary.
Default: ~/.agenterm/config.toml. Override with AGENTERM_CONFIG or --config.
No affiliation. Independent project.
cmd/agenterm/ CLI entry
internal/config/ TOML + env + provider presets
internal/llm/ OpenAI-compatible client (stream + tools)
internal/agent/ Multi-turn tool loop
internal/tools/ Built-in tools
internal/mcp/ MCP client
internal/tui/ Bubble Tea UI
scripts/install.sh End-user installer
| Piece | Tech |
|---|---|
| Language | Go |
| CLI | Cobra |
| TUI | Bubble Tea · Bubbles · Lip Gloss |
| Markdown | Glamour |
| MCP | official Go MCP SDK |
Chat replies come from the LLM. MCP only supplies tools. Models are not embedded in the binary.
More detail: docs/how-it-works.md.
Roadmap (Grok-class UX): docs/grok-parity-roadmap.md.
Notes: Ollama and SGLang must expose the OpenAI-compatible API (default for both). Tool quality depends on the model (qwen2.5, llama3.1+ work better than tiny models).
On version tags, GitHub Actions publishes multi-platform binaries and a GHCR image:
| Asset | Platforms |
|---|---|
agenterm-linux-amd64 |
Linux x86_64 |
agenterm-linux-arm64 |
Linux aarch64 |
agenterm-darwin-amd64 |
macOS Intel |
agenterm-darwin-arm64 |
macOS Apple Silicon |
agenterm-windows-amd64.exe |
Windows |
ghcr.io/saurabhahuja71/agenterm:vX.Y.Z |
Container |
git tag v0.2.1 && git push origin v0.2.1
make dist VERSION=0.2.1 # local cross-build| # | Lab | Focus |
|---|---|---|
| 0 | mcp-demo | Build MCP server tools |
| 1 (this) | agenterm | Terminal agent + MCP client |
| 2 | agentic-ai-sample | LangChain ReAct research agent |
| 3 | workshop-redis-ai | Redis vector search + RAG workshop |
Hub: learning-path
ai-agent ollama mcp golang tui openai xai terminal llm coding-agent tutorial education
MIT — free to use, modify, and distribute.
Suggested topics for the About sidebar:
ai · terminal · cli · tui · ollama · sglang · llm · openai · mcp · agent · golang · bubbletea · coding-assistant · local-llm