How to Get the Most Out of a 24GB or 32GB VRAM GPU for Private Local AI ModelsLearn how to run private local AI models efficiently on a 24GB or 32GB VRAM GPU. This tutorial covers Ollama setup, model selection, quantization, context size, VRAM monitoring, testing, and troubleshooting for real-world local LLM workflows.

Table of Contents

Introduction: What You Will Build

A 24GB or 32GB VRAM graphics card is one of the best price-to-performance tiers for running private local AI models. With the right setup, you can run strong 7B, 8B, 14B, and even some 30B to 32B models locally for coding, writing, document analysis, research, and private chat without sending your data to a cloud provider.

In this tutorial, you will set up a practical local LLM workstation, choose the right model size and quantization level, configure context length correctly, monitor VRAM usage, and test whether your GPU is being used efficiently. The goal is not just to load a model. The goal is to have enough usable context for real work while keeping the system stable and responsive.

You will learn how to:

  • Install and verify GPU support.
  • Run private local models with Ollama.
  • Choose suitable models for 24GB and 32GB VRAM cards.
  • Set context size without wasting VRAM.
  • Understand quantization, KV cache, and GPU offload.
  • Test performance and troubleshoot common memory errors.

Important principle: Do not spend all your VRAM on model weights. Leave room for context, runtime overhead, and the KV cache. A model that loads successfully can still crash or slow down badly when you increase context too much.

Prerequisites: What You Need Before You Start

Before you begin, prepare the following hardware and software.

Hardware

  • A GPU with 24GB or 32GB VRAM, such as an RTX 3090, RTX 4090, RTX 6000-class card, or similar.
  • At least 32GB system RAM. Use 64GB or more if you plan to offload larger models to RAM.
  • Fast SSD storage with at least 100GB free. Local models can be large.
  • A stable power supply and good case airflow.

Software

  • Windows 11, Ubuntu Linux, or another modern Linux distribution.
  • Latest NVIDIA GPU driver if you use an NVIDIA card.
  • Ollama for the beginner-friendly setup.
  • Optional: Docker and Open WebUI if you want a browser-based ChatGPT-like interface.
  • Optional: llama.cpp if you want deeper control over GGUF models and GPU offload.

Recommended Knowledge

  • Basic terminal or PowerShell usage.
  • Basic understanding of files, folders, and command-line commands.
  • Basic awareness that larger AI models need more VRAM and run slower.

Step 1: Verify Your GPU and Driver

Start by confirming that your GPU is detected correctly. This step is necessary because local AI performance depends heavily on GPU acceleration. If your GPU driver is missing or broken, the model may run on CPU and feel unusably slow.

Run This Command

nvidia-smi

Expected Result

You should see a table showing your GPU name, driver version, CUDA version, temperature, power usage, and memory usage.

+---------------------------------------------------------------------------------------+
| NVIDIA-SMI              Driver Version: 555.xx        CUDA Version: 12.x              |
| GPU  Name               Memory-Usage                                                  |
| 0    NVIDIA RTX 4090    500MiB / 24564MiB                                             |
+---------------------------------------------------------------------------------------+

If you see your card and the correct VRAM amount, continue. For a 24GB card, expect around 24564 MiB. For a 32GB card, expect around 32768 MiB, depending on the model and driver reporting.

Common Error

If the command says nvidia-smi not found, install or reinstall your NVIDIA driver. On Ubuntu, use your distribution driver manager or run:

sudo ubuntu-drivers autoinstall
sudo reboot

Warning: Reboot after installing GPU drivers. Do not start troubleshooting AI tools until the driver is working correctly.

Step 2: Understand the VRAM Budget Before Choosing a Model

Do this before downloading large models. Your VRAM is used by more than the model itself. A local LLM session consumes memory in several places:

  • Model weights: The actual neural network parameters.
  • KV cache: Memory used to store attention information for the current context.
  • Runtime overhead: Memory needed by Ollama, llama.cpp, CUDA, drivers, and buffers.
  • Desktop overhead: Memory used by browsers, video players, games, and other GPU apps.

Leave 20% to 30% of VRAM free whenever possible. This headroom prevents crashes when you increase context, paste long documents, or use a web UI with longer conversations.

Practical VRAM Rule of Thumb

GPU VRAMComfortable Model RangeBest Default ContextNotes
24GB7B to 14B easily, some 30B to 32B quantized4K to 8KOptimize aggressively and avoid max context by default.
32GB7B to 14B comfortably, 30B to 32B more practical8K to 16KMore room for context, better quantization, or larger models.

For actual work, a fast 7B to 14B model with 8K context is often more useful than a huge model that runs slowly with a tiny context window. Choose the smallest model that solves the task well.

Step 3: Install Ollama for a Simple Local AI Setup

Ollama is the easiest way to start running local models privately. It handles model downloads, serving, GPU usage, and a simple command-line chat interface. Use it first, even if you later move to llama.cpp or vLLM.

Install on Linux

curl -fsSL https://ollama.com/install.sh | sh

Install on Windows or macOS

Download the installer from the Ollama website, install it, and open a new terminal or PowerShell window.

Verify the Installation

ollama --version

Expected Result

ollama version 0.x.x

If the command works, Ollama is installed. If it does not work, close and reopen your terminal. On Windows, make sure Ollama is running in the background from the system tray.

Step 4: Start With a Small Model to Confirm GPU Acceleration

Do not begin with a giant model. First, run a smaller model to verify that the system works and your GPU is being used. This saves time and makes troubleshooting easier.

Run a Good General-Purpose Model

ollama run llama3.1:8b

Then type:

Explain the difference between VRAM and system RAM in five bullet points.

Monitor GPU Usage

Open a second terminal and run:

watch -n 1 nvidia-smi

On Windows PowerShell, run:

nvidia-smi -l 1

Expected Result

When the model is generating text, GPU memory usage should increase and GPU utilization should rise. If VRAM usage does not change and generation is extremely slow, the model may be running on CPU.

Screenshot description: At this stage, a screenshot should show a terminal running Ollama on the left and nvidia-smi on the right. The GPU memory column should show several gigabytes in use while the model responds.

Step 5: Choose the Right Model for Your 24GB or 32GB Card

Now choose a model based on your actual work. Model selection matters more than chasing the largest parameter count. A coding model is better for coding. A general chat model is better for writing, planning, summarization, and reasoning. A long-context model is better for document analysis.

Recommended Starting Models

Use CaseRecommended SizeExamplesWhy It Works
General assistant7B to 14BLlama, Mistral, QwenFast, useful, and easy to fit with context.
Coding7B to 14B, then 30B if neededQwen Coder, DeepSeek CoderSpecialized training improves code output.
Writing and editing8B to 14BLlama, Mistral, Gemma-style modelsGood quality without wasting VRAM.
Document Q and A8B to 14B with 8K contextLong-context instruct modelsContext length matters more than raw size.
Advanced reasoning14B to 32BLarger instruct modelsBetter reasoning, but slower and more memory-heavy.

Good Defaults for 24GB VRAM

  • Start with an 8B model.
  • Use 4-bit quantization by default.
  • Set context to 4096 or 8192 tokens.
  • Try 14B after confirming stable performance.
  • Try 30B to 32B only if you accept lower speed or tighter context.

Good Defaults for 32GB VRAM

  • Start with 8B or 14B for speed.
  • Use 4-bit or 5-bit quantization depending on availability.
  • Set context to 8192 tokens first.
  • Increase to 16384 tokens if VRAM remains available.
  • Use 30B to 32B models when quality matters more than speed.

Step 6: Set Context Size Deliberately

Context size is one of the most misunderstood settings in local AI. It controls how much text the model can consider at once. This includes your prompt, prior chat history, retrieved documents, code snippets, and the model response.

Larger context sounds better, but it uses more VRAM through the KV cache. If you set context too high, you may waste memory, reduce speed, or crash with out-of-memory errors.

Recommended Context Sizes

TaskUseful ContextWhy
Quick chat4096 tokensEnough for short instructions and answers.
Writing and editing4096 to 8192 tokensEnough for article sections, outlines, and revisions.
Coding one file8192 tokensEnough for a moderate source file plus instructions.
Coding multiple files8192 to 16384 tokensUseful when pasting related files or error logs.
Document Q and A8192 to 16384 tokensUseful for retrieved chunks and summaries.
Long research sessions16384 tokens or moreOnly use if your model and VRAM can handle it.

For a 24GB GPU, use 4096 to 8192 tokens for most work. For a 32GB GPU, use 8192 tokens as your everyday default and increase to 16384 when needed.

Step 7: Create an Ollama Model With a Custom Context Window

Ollama lets you create a custom model configuration with a specific context size. This is useful because you can keep multiple versions of the same model: one fast 4K version, one balanced 8K version, and one long-context version.

Create an 8K Context Model

Create a file named Modelfile:

cat > Modelfile <<'EOF'
FROM llama3.1:8b
PARAMETER num_ctx 8192
PARAMETER temperature 0.7
PARAMETER top_p 0.9
EOF

Build the custom model:

ollama create llama31-8b-8k -f Modelfile

Run it:

ollama run llama31-8b-8k

Expected Result

Ollama should create a new local model entry and start a chat session using an 8192-token context window.

success
>>> Send a message

Why This Step Matters

Setting context explicitly prevents accidental overuse of VRAM. It also helps you compare model behavior fairly. If one model runs at 4K and another runs at 16K, you are not comparing only model quality; you are also comparing memory pressure.

Step 8: Use Quantization Correctly

Quantization reduces model memory usage by storing weights with fewer bits. This is the key technique that makes local LLMs practical on consumer GPUs.

Common Quantization Levels

QuantizationMemory UseQualityBest For
FP16 or BF16Very highHighestSmall models or professional GPUs.
8-bitMediumVery goodWhen you have enough VRAM and want quality.
5-bitLowerGood32GB cards and balanced quality.
4-bitLowGood enough for most useDefault choice for 24GB and many 32GB setups.
3-bit or lowerVery lowNoticeable quality lossOnly when fitting a model matters more than quality.

For 24GB VRAM, use 4-bit as the default. For 32GB VRAM, use 4-bit for larger models and consider 5-bit for smaller models if speed and VRAM remain acceptable.

Warning: Do not assume a larger, heavily quantized model is always better than a smaller, cleaner model. A strong 14B model at a good quantization level can outperform a larger model that has been compressed too aggressively.

Step 9: Install Open WebUI for a Browser-Based Private Chat Interface

Ollama works well in the terminal, but a browser interface is easier for daily work. Open WebUI gives you a private ChatGPT-like interface that connects to your local Ollama server.

Install Docker First

Verify Docker works:

docker --version

Run Open WebUI

docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data --name open-webui --restart always ghcr.io/open-webui/open-webui:main

Open the Interface

Go to:

http://localhost:3000

Expected Result

You should see a web interface where you can create a local account and choose your Ollama models from a dropdown.

Screenshot description: The page should show a clean chat layout with a model selector at the top, a message box at the bottom, and previous conversations in a sidebar.

Privacy note: Keep this service bound to your local machine unless you know how to secure it. If you expose it to your network or the internet, add authentication, firewall rules, and HTTPS.

Step 10: Monitor VRAM While You Increase Context

After your model works, increase context slowly and watch VRAM. This is the safest way to find your practical limit.

Test Sequence

  1. Run the model at 4096 context.
  2. Ask a short question.
  3. Watch VRAM usage with nvidia-smi.
  4. Create an 8192 context version.
  5. Paste a longer prompt or code file.
  6. Watch VRAM again.
  7. Only move to 16384 if you still have several GB free.

Example Long Prompt Test

Analyze the following code. Identify bugs, security issues, and performance problems. Then rewrite the code with comments explaining each change.

Paste your code here.

Expected Result

VRAM usage should increase as prompts get longer and context grows. Generation may slow down with longer context. This is normal.

If VRAM approaches the card limit, reduce context or use a smaller model. Do not run your card constantly at the absolute limit if you want a stable workstation.

Step 11: Use the Right Settings for Real Work

Use these settings as practical defaults. Adjust them only after you have a stable baseline.

Must-Have Settings

  • Context: 4096 to 8192 for 24GB, 8192 to 16384 for 32GB.
  • Quantization: 4-bit for most models on 24GB, 4-bit or 5-bit on 32GB.
  • Temperature: 0.2 to 0.4 for coding and factual work, 0.7 to 0.9 for brainstorming and writing.
  • Top-p: 0.9 is a good general default.
  • Model choice: Use specialized coding models for code and general instruct models for chat or writing.
  • VRAM headroom: Keep several GB free if possible.

Example Coding Configuration

cat > Modelfile <<'EOF'
FROM qwen2.5-coder:14b
PARAMETER num_ctx 8192
PARAMETER temperature 0.2
PARAMETER top_p 0.9
EOF

ollama create coder-14b-8k -f Modelfile
ollama run coder-14b-8k

Example Writing Configuration

cat > Modelfile <<'EOF'
FROM llama3.1:8b
PARAMETER num_ctx 8192
PARAMETER temperature 0.8
PARAMETER top_p 0.9
EOF

ollama create writer-8b-8k -f Modelfile
ollama run writer-8b-8k

Step 12: Build a Practical Model Lineup

Instead of using one model for everything, create a small local model lineup. This gives you better results and avoids wasting VRAM.

Recommended Lineup for 24GB VRAM

  • Fast assistant: 7B or 8B model with 4096 or 8192 context.
  • Coding assistant: 7B to 14B coding model with 8192 context.
  • Quality model: 14B model or carefully quantized 30B model with reduced context.
  • Document model: 8B model with 8192 context for summaries and Q and A.

Recommended Lineup for 32GB VRAM

  • Fast assistant: 8B model with 8192 context.
  • Coding assistant: 14B coding model with 8192 or 16384 context.
  • Quality model: 30B to 32B quantized model with 8192 context.
  • Document model: Long-context 8B or 14B model with 16384 context.

This approach is better than trying to force one huge model to handle every task. Use fast models for routine work and larger models only when quality matters.

Testing: Verify That Your Local AI Setup Works

Run these tests after setup. They confirm that your model loads, uses the GPU, handles context, and produces useful output.

Test 1: GPU Usage

nvidia-smi -l 1

Run a model in another terminal. Expected result: VRAM usage increases and GPU utilization rises while text is generated.

Test 2: Context Handling

Paste a 2,000 to 4,000 word document and ask:

Summarize this document into 10 bullet points. Then list the three most important action items.

Expected result: The model should reference information from the pasted text instead of giving a generic answer.

Test 3: Coding Quality

Paste a real function from your project and ask:

Review this function for bugs. Explain each issue and provide a corrected version.

Expected result: A coding model should identify specific logic, security, or style issues and produce usable code.

Test 4: Stability

Run a long chat for 10 to 15 minutes. Watch for crashes, out-of-memory errors, or severe slowdowns. If the model fails, reduce context first. If it still fails, use a smaller or more aggressively quantized model.

Troubleshooting: Common Issues and Fixes

Problem: The Model Runs Very Slowly

Likely causes: The model is running on CPU, too many layers are offloaded to RAM, or the model is too large.

Fix: Check nvidia-smi. If GPU usage is low, reinstall GPU drivers or try a smaller model. Start with an 8B model to confirm acceleration.

Problem: CUDA Out of Memory

Likely causes: Context is too large, the model is too large, or other apps are using VRAM.

Fix: Close GPU-heavy apps, reduce context from 16384 to 8192 or 4096, or use a smaller quantized model.

Problem: The Model Loads But Crashes During Long Prompts

Likely cause: The model weights fit, but the KV cache does not have enough room as context grows.

Fix: Lower num_ctx. This is the first setting to reduce when long prompts fail.

Problem: Answers Get Worse With a Huge Model

Likely causes: The model is too heavily quantized, too slow, or not suited to the task.

Fix: Try a smaller, higher-quality model. For coding, use a coding-specific model. For writing, use a strong instruct model.

Problem: Open WebUI Cannot See Ollama Models

Likely causes: Ollama is not running, Docker cannot reach the host, or the container was started without the correct host mapping.

Fix: Restart Ollama, then recreate the Open WebUI container with the command shown earlier.

Problem: Your Desktop Becomes Laggy

Likely causes: VRAM is full, the GPU is under sustained load, or your browser is also consuming GPU memory.

Fix: Reduce context, close browser tabs, disable hardware acceleration in nonessential apps, or use a smaller model for daily chat.

Next Steps: Extend Your Private Local AI Workstation

Once your basic setup is stable, improve it in these ways.

1. Add Retrieval-Augmented Generation

Use a document search or RAG system to retrieve only the most relevant chunks instead of pasting entire documents into context. This saves VRAM and improves answer quality.

2. Add Project-Specific Coding Workflows

Create prompts for code review, test generation, refactoring, and documentation. Use an 8K or 16K context coding model when working with multiple files.

3. Try llama.cpp for More Control

If you want deeper control over GGUF files, GPU layers, batch size, and CPU/GPU split, try llama.cpp. It is excellent for advanced tuning and hybrid offloading.

4. Consider vLLM for Multi-User Serving

If you want to serve local models to a team or multiple applications, consider a server-oriented runtime. Throughput and concurrency settings become more important than single-user chat speed.

5. Create Separate Profiles

Create one model profile for speed, one for coding, one for long documents, and one for high-quality reasoning. Switch profiles instead of constantly changing settings.

Conclusion

A 24GB or 32GB VRAM GPU is powerful enough for serious private local AI work if you configure it correctly. The most important lesson is to balance model size, quantization, and context length. Do not simply load the biggest model you can find. Use 4-bit quantization as your default, leave VRAM headroom, and increase context only when the task needs it.

For 24GB cards, focus on 7B to 14B models with 4096 to 8192 context for daily use. Try larger 30B-class models only when you are comfortable tuning memory and accepting slower output. For 32GB cards, use 8B and 14B models comfortably with 8192 to 16384 context, and keep 30B to 32B models available for higher-quality tasks.

The best local AI setup is not the largest possible model. It is the setup that answers accurately, runs privately, stays stable, and gives you enough context to do real work.

Leave a Reply