Kagi Assistant
Kagi Assistant combines the top large language models (LLMs) with optional results from Kagi Search, making it the perfect companion for creative, research, and programming tasks — alongside everything else you can think of!
All this is included in a single subscription!
Features
- Access to the latest and most performant large language models from OpenAI, Anthropic, Meta, Google, Mistral, Amazon, Alibaba, and DeepSeek
- Multiple custom assistants
- The ability to control whether the Assistant has web access (powered by Kagi Search)
- Applying Kagi Search Lenses and Personalized Results to the Assistant searches
- Saving Assistant threads
- Uploading files to use as context
- Altering the Assistant configuration within the thread
- For example, you can ask the initial question with web access enabled and then disable it for subsequent questions
- It is also possible to switch to a different LLM in the middle of a thread
- Code syntax highlighting
- Keyboard Shortcuts
- Export conversations as markdown or JSON
- Share threads with others using a link
- Voice input
Privacy
When you use the Assistant by Kagi, your data is never used to train AI models (not by us or by the LLM providers), and no account information is shared with the LLM providers. By default, threads are deleted after 24 hours of inactivity. This behavior can be adjusted in the settings.
Using the Assistant
Kagi Assistant can be accessed via the apps menu located in the top right corner of all Kagi pages or by using bangs in search. You can also use this direct link.
When you first access the Assistant, you will be greeted by a familiar-looking landing page, allowing you to get right into using it. You can either type your prompt or use voice input by pressing the microphone symbol. You can choose which LLM you wish to use by opening the dropdown menu just below the prompt field.
The Assistant's web access can be toggled via the button below the prompt field.
Which model to choose
There is no definite answer to the question of what the best LLM is. As the number of competing models increases, users may find it difficult to find the right one for their task. To aid in this, Kagi maintains a list of recommended models at the top of the LLM list.

Kagi recommended models as of July 1, 2026.
The recommendations are based on the Kagi LLM Benchmarking Project. The benchmark tests measure model quality in various scenarios.
Another important aspect is the privacy policy of the model provider. See our LLM Privacy Comparison for a detailed overview of how each provider handles your data.
Settings
You can manage your Assistant settings by clicking the Settings button in the bottom-left corner of the Assistant window. Depending on device type and window size, you may have to first open the sidebar by clicking the sidebar icon on desktop or on mobile devices.
General
Thread Saving
- This allows you to configure the thread retention setting. The options are temporary (threads expire 24 hours after the last message) or permanent.
Default Assistant
- Choose which assistant is selected by default when opening the Assistant.
Custom Instructions
- This text box will allow you to provide instructions and context to the Assistant. These instructions are applied to all interactions unless you are using a custom assistant with its own instructions.
Appearance
Workspace area
- Configure whether the assistant chat utilizes the entire screen (Wide) or is displayed in a more compact view (Standard).
Font Size
- Set the font size for the entire Assistant site (small, medium, normal, large, larger).
Theme
- Choose the theme for the Assistant.
- Note that this setting is saved locally on your device. If you have a browser extension that clears local storage, it will also reset this setting.
Custom Assistants
- Kagi Assistant supports creating custom assistants. This allows you to have a customized assistant with the desired LLM, custom instructions, and web search access (including lenses and/or personalized results).
- Please see Custom Assistants for further information.
Shortcuts
You can configure the keyboard shortcuts to fit your workflow.
Prompt submission trigger controls whether the shortcut to send a message is Enter or ⌘/Ctrl + Enter. Adding a modifier key makes it less likely to accidentally submit a prompt before you are finished writing it.
Below you can find the default keybinds.
| Mac Shortcut | PC shortcut | Action |
|---|---|---|
| / | / | Focus prompt |
| ⌘ + K | Ctrl + K | New thread |
| ⌘ + Shift + Backspace | Ctrl + Shift + Backspace | Delete thread |
| ⌘ + Shift + S | Ctrl + Shift + S | Toggle sidebar |
| ⌘ + Shift + C | Ctrl + Shift + C | Copy last response |
| ⌘ + Shift + E | Ctrl + Shift + E | Edit last message |
| ⌘ + Shift + G | Ctrl + Shift + G | Regenerate last message |
| ⌘ + Shift + M | Ctrl + Shift + M | Open model selector |
| ⌘ + U | Ctrl + U | Upload file |
| ⌘ + . | Ctrl + . | Show keyboard shortcuts |
Threads
Interactions with the Assistant are stored in threads.
The search bar enables you to find that one elusive thread.
By default, threads are kept for 24 hours after the last message. If keeping threads alive permanently better fits your workflow, you can adjust this in the settings.
Please note that the thread-saving setting is applied when the thread is created.
Threads can be renamed, pinned, made permanent, shared, added to folders, exported, and deleted via the ⋮ button, which is displayed when you hover over the thread.
Folders
Folders are how Kagi Assistant enables you to organize your threads. By utilizing folders, you can group threads by topic, so your work and leisure are never mixed.
Folders are mutually exclusive, so a single thread can only be in one folder at a time.
You can customize the look of your folders in the sidebar, making them easier to distinguish.
Uploading Files to Assistant
Kagi Assistant supports file uploads, allowing you to provide additional context or information for your queries.
This can be useful for tasks like:
- Summarizing a document
- Extracting key insights from a report
- Analyzing data in a spreadsheet
- Describing an image
- Distilling main points from an audio file
To upload a file:
- Click the paperclip icon in the prompt input box or use the upload shortcut (⌘/Ctrl + U by default).
- Select the file or image you wish to upload.
- Provide a prompt with instructions to process the file, or leave it blank to summarize it.
Important considerations for file uploads:
- File size limit: The maximum file size for uploads is 30 MB.
- Processing time: Larger files may take a few moments to process.
- Context retention: Uploaded file content remains in the conversation context for subsequent messages.
The Assistant supports various file formats across different categories, including:
| File Type | Supported Formats |
|---|---|
| Text | txt, text, md (and other text-based formats) |
| Rich Format | pdf, docx, pptx |
| Spreadsheets | csv, tsv, xlsx, json, jsonl |
| Image | jpg, jpeg, png, gif, tiff, tif, webp |
| Audio | 3gpp, aa, aac, aax, act, aiff, amr, ape, au, awb, dct, dss, dvf, flac, gsm, iklax, ivs, m4a, m4b, m4p, mp4, mmf, mp3, mpc, msv, ogg, opus, ra, rm, sln, tta, vox, wav, wma, wvpla |
Note: Unsupported formats may be treated as binary files.
Fetching online content
Assistant can fetch webpages and online documents (up to 50 MB) to use them as context for your conversation. To use this feature, simply paste the URL in your Assistant conversation (make sure the Entire Web toggle is on).
Available LLMs
| Developer | Model | Profile | Plan |
|---|---|---|---|
| Alibaba | Qwen3-Coder | qwen-3-coder | All |
| Alibaba | Qwen3.7 Plus | qwen-3-7-plus | All |
| Anthropic | Claude 4.5 Haiku | claude-4-haiku | Ultimate |
| Anthropic | Claude 4.6 Sonnet | claude-4-sonnet | Ultimate |
| Anthropic | Claude Opus 5 | claude-5-opus | Ultimate |
| Anthropic | Claude Fable 5 | claude-5-fable | Ultimate |
| Anthropic | Claude 4.6 Sonnet (Reasoning) | claude-4-sonnet-thinking | Ultimate |
| Anthropic | Claude Opus 5 (Reasoning) | claude-5-opus-thinking | Ultimate |
| Anthropic | Claude Fable 5 (Reasoning) | claude-5-fable-thinking | Ultimate |
| Anthropic | Claude 4.5 Haiku (Reasoning) | claude-4-haiku-thinking | Ultimate |
| DeepSeek | DeepSeek V4 Pro | deepseek-v4-pro | Ultimate |
| DeepSeek | DeepSeek V4 Flash | deepseek-v4-flash | All |
| Gemini 3.6 Flash | gemini-3-6-flash | Ultimate | |
| Gemini 3.5 Flash Lite | gemini-3-5-flash-lite | Ultimate | |
| Gemini 2.5 Pro | gemini-2-5-pro | Ultimate | |
| Gemini 3.1 Pro (Preview) | gemini-3-pro | Ultimate | |
| Gemini 3.1 Flash Lite | gemini-3-1-flash-lite | All | |
| Gemma 4 31B | gemma-4-31b | All | |
| Mistral AI | Mistral Small 4 | mistral-small-4 | All |
| Mistral AI | Mistral Small 3 | mistral-small | All |
| Mistral AI | Mistral Medium 3.5 | mistral-medium-3-5 | Ultimate |
| Mistral AI | Mistral Large 3 | mistral-large | All |
| MiniMax | MiniMax M3 | minimax-m3 | All |
| Moonshot AI | Kimi K2.7 Code | kimi-k2-7-code | Ultimate |
| Nous Research | Hermes-4-405B (reasoning) | hermes-4-405b-thinking | All |
| OpenAI | GPT 5.4 Mini | gpt-5-4-mini | All |
| OpenAI | GPT 5.4 Nano | gpt-5-4-nano | All |
| OpenAI | GPT OSS 120B | gpt-oss-120b | All |
| OpenAI | GPT 5.5 | gpt-5-5 | Ultimate |
| OpenAI | GPT 5.6 Sol | gpt-5-6-sol | Ultimate |
| OpenAI | GPT 5.6 Terra | gpt-5-6-terra | Ultimate |
| OpenAI | GPT 5.6 Luna | gpt-5-6-luna | Ultimate |
| OpenAI | ChatGPT | chatgpt-4o | Ultimate |
| xAI | Grok 4.3 | grok-4-20 | All |
| xAI | Grok 4.5 | grok-4-5 | Ultimate |
| Z.ai | GLM-5.2 | glm-5-2 | Ultimate |
| Z.ai | GLM-5.2 (reasoning) | glm-5-2-thinking | Ultimate |
| Z.ai | GLM-4.7 | glm-4-7 | All |
| Z.ai | GLM-4.7 (reasoning) | glm-4-7-thinking | All |
You can learn more about how these models compare in the Kagi LLM Benchmarking Project page.
For more information about each model and its privacy practices, including details about providers, see our LLM Privacy page.
Bangs
You can quickly access Assistant using the following bangs:
!ai,!as,!assistant,!research,!answer,!discuss,!expert,!llm,!custom, and!asst: These bangs direct you to the general Assistant interface for various types of queries.!chat: This bang accesses Assistant with internet access turned off.!code: Use this bang to access the built-in Code Custom Assistant, which is tailored for coding-related queries.!kiand!quick: These bangs access Assistant with the Quick profile, providing a fast, direct answer to your queries.!study: This bang opens the Assistant with the Study assistant.!news: This bang opens the Assistant with the News custom assistant.
See Custom Assistants for further information about the custom assistants.
Each bang is designed to optimize your search experience by directing you to the most appropriate version of Assistant for your needs.
URL Parameters
You can specify a particular model in the Assistant's URL by including a profile parameter. https://kagi.com/assistant?profile=gpt-5 The available model names can be found in the table above.
The internet parameter can be used to turn on and off internet access, set to true to enable, anything else to disable. This overrides the internet setting of the profile used.
The lens parameter can be used to set the lens if internet access is enabled. The value of this is the lowercase format of the lens name, for example, https://kagi.com/assistant?lens=programming will use the Programming lens.
The q parameter can be used to submit a prompt immediately after the page loads. The qvalue parameter can be used to prefill the prompt box without submitting it.
Here is an example of a URL that enables internet access, uses the Claude 4 Sonnet model, applies the Recipes lens, and submits a prompt immediately. You might use it as a target for a custom bang. https://kagi.com/assistant?profile=claude-4-sonnet&internet=true&lens=recipes&q=%s
Availability
Assistant is available to all members. However, premium models are only available in our Ultimate plan. If you are on a different plan and you need access to these models, you can upgrade from the Billing Settings page.
We also offer an Ultimate upgrade for Family Plans. You can upgrade from the Family Management page.
Usage Limits
Context window limit
There's no fixed limit on conversation length. We automatically optimize lengthy chats behind the scenes to maintain performance using summarization and truncation techniques. Currently, these techniques are used when the chat history reaches about 32k tokens in size.
Input limitations
Text input
- Maximum 100,000 characters per message
- Text exceeding this limit will be automatically truncated
File uploads
- Maximum total size: 30 MB (applies to single or multiple files)
- URL content: 50 MB maximum retrievable size
Custom Instructions
- Maximum 20,000 characters for custom Assistant instructions
Fair Use Policy
We use a value-based usage system to maintain high-quality service for all users:
- Your monthly plan determines your token usage allowance.
- For example, a $25 monthly plan provides up to $25 worth of token usage across all models.
- For yearly plans, you get access to the full year's worth of token usage at the start of the plan.
- For instance, the Ultimate yearly plan allows up to $270 worth of token usage for the entire year.
- A 20% margin markup is included in token usage cost calculations to cover search queries, infrastructure, and development costs.
- For example, $25 token usage consists of $20 for raw token costs and $5 for operational costs.
- Users will receive an in-app reminder as they near their usage limit. If the limit is exceeded, new AI interactions will be disabled until they either renew their plan early or the next billing cycle begins.
- Note: We will soon introduce the option to purchase top-up credits, allowing you to extend Assistant usage beyond fair-use limits with an amount of your choice. These credits can then also be used for other Kagi products such as the API.
For additional questions about these limitations or policies, please contact our support team.
Tips to reduce token usage
Here are some suggestions to reduce token usage:
- Use less expensive models for simple tasks like summarization or basic information extraction. Our LLM Benchmarking project page contains cost information for the different models.
- Create new threads for unrelated questions rather than continuing in the same conversation.
- Be specific and concise in your prompts to get more focused responses.
- Use the "Edit Prompt” feature (pencil icon) to refine your question instead of sending multiple clarifications.
- Disable web access when you don't need internet information.
- Limit file uploads to only what's necessary for your query.
- Break complex tasks into smaller, focused questions across multiple threads.
- Use custom instructions to request consistently concise responses.
- Leverage specialized custom assistants optimized for specific tasks.
- Download and delete completed threads to avoid accidentally continuing old conversations.
FAQ
Q: What is Kagi’s stance about using LLMs in search?
A: We continue to relentlessly focus on the core search experience and build thoughtfully integrated features on top of it. Read more about it in our AI Integration Philosophy page.
Q: Why is Assistant unable to access the web?
A: Check to make sure you haven't disabled web access. Look for the lens selector dropdown under the prompt entry box to the right of the model selector. Tap the globe to enable or disable web access. It should be colored light purple when web access is enabled. If you are still having issues, please contact support.