Gemini capabilities
19 mapped capabilities, each graded and dated. The map shows what Gemini can do; the audit shows whether it’s worth consolidating — and a guide shows how to move.
Capabilities
Canvas
provisionalverified ~2 months agoCanvas is Gemini's side-by-side interactive workspace for collaborating with the model on documents, slide presentations, code, and runnable apps without leaving the chat. The Canvas panel opens to the right of the conversation so users can iterate on a draft, deck, or prototype while still prompting in the main thread.
Code execution
provisionalverified ~2 months agoGemini executes Python (and renders web apps) inside the Canvas code-interpreter sandbox to compute results, plot data, and prototype small applications. Code in Canvas is auto-saved; runnable previews appear in the Canvas panel; Python can be one-click exported to a Google Colab notebook for further iteration.
Context grouping (Projects)
canonicalverified ~2 months agoGemini's project-grouping feature shipped under the name 'notebooks' (announced on blog.google 2026-04-08). A notebook is a dedicated project space that holds related chats, source files, and per-notebook custom instructions, providing 'a continuous chat experience that remembers your sources, instructions, and ongoing discussions'. Notebooks sync bidirectionally with NotebookLM — sources added in one appear in the other.
Custom assistants (Gems)
provisionalverified ~2 months agoGems are Google's equivalent of GPTs: persistent custom AI assistants with their own name, instructions, and optional knowledge files. Users can create their own Gems, use Google's suite of premade Gems (Brainstormer, Career guide, Coding partner, Learning coach, Writing editor), and share a custom Gem with specific people (Drive-style link/email sharing, web only, since Sept 2025). NOTE: there is no public consumer Gem gallery/marketplace for browsing strangers' Gems - the official overview page lists only premade Gems plus build-your-own; the discovery 'Agent Gallery' is a Gemini Enterprise feature, not the consumer app.
Deep Research
provisionalverified ~2 months agoDeep Research is Gemini's agentic web-research mode: the user states a topic, Gemini generates a research plan, autonomously browses dozens to hundreds of web pages (optionally Gmail/Drive/Chat), iterates, and produces a long-form structured report with inline citations. On the Ultra plan, app reports can include visuals — charts, diagrams, and interactive simulators; a higher-tier 'Deep Research Max' uses extended test-time compute for the highest-quality reports (Gemini API).
File & document handling
canonicalverified ~2 months agoGemini accepts a wide range of file types including documents, images, audio, video, code folders, and GitHub repositories. Per-prompt and per-file size limits vary by plan; larger context window and longer media on paid plans.
Gemini Agent / Spark (agentic action)
provisionalverified 25 days agoGemini Agent and Gemini Spark are the consumer-facing agentic capabilities: Agent performs multi-step web tasks on the user's behalf (booking, filling forms, comparing options) inside the Gemini app, and Spark (announced at I/O 2026) is a 24/7 personal agent that runs proactively in the background, automating workflows and managing schedules for ongoing tasks. Both are gated to Google AI Ultra; Spark is the headline agentic feature of the post-I/O-2026 Ultra tiers ($99.99 and $199.99), US-only.
Gemini Live
provisionalverified ~2 months agoGemini Live is a real-time, hands-free spoken conversation mode where the user can talk to Gemini, interrupt, change topics, and optionally share camera or screen so Gemini can see what the user is looking at and respond verbally. Available primarily on mobile (Android, iOS) with limited Live access on the web.
Gemini in Workspace (Gmail, Docs, Slides, Sheets, Meet, Vids)
provisionalverified ~2 months agoBeyond the standalone gemini.google.com app, Gemini is embedded into Google Workspace apps as a side-panel assistant: Help me write in Gmail and Docs, Help me visualize in Slides, formula/insights in Sheets, Take notes for me in Meet, and AI video drafting in Vids. Workspace embeds inherit Workspace data-handling guarantees (no training on user data).
Image generation
provisionalverified ~2 months agoGemini generates and edits images natively via the Nano Banana family (Nano Banana 2 / Gemini 3.1 Flash Image as the default, Nano Banana Pro built on Gemini 3 Pro for higher quality) plus Imagen-derived models exposed via Google Flow. Supports text-to-image, image editing, character/scene consistency, and personalized images using Google Photos faces with consent.
Integrations / connectors (Connected Apps / Workspace apps)
provisionalverified 25 days agoGemini integrates with Google Workspace (Gmail, Drive, Docs, Calendar, Keep, Tasks) and other Google services (Maps, YouTube, Flights, Hotels). Formerly invoked as '@-extensions', many are now direct integrations.
Jules (async coding agent)
provisionalverified ~2 months agoJules is Google's asynchronous, autonomous coding agent (from Google Labs, powered by Gemini). You hand it a task instead of pair-programming: it clones your repository into a secure Google Cloud VM, reads the full project context, writes a plan, makes multi-file changes (bug fixes, tests, feature work, dependency bumps), and opens a GitHub pull request for you to review and merge while you do other work. It positions against OpenAI Codex and Devin as a background agent rather than an in-editor copilot.
Memory / conversation history (Gemini Apps Activity)
provisionalverified ~2 months agoGemini retains full chat history under the 'Keep Activity' setting (the current name for what was 'Gemini Apps Activity') with chronological prompts and responses. Gemini can also recall past chats to personalize responses in new conversations via the separate 'Memory' control under Personal Intelligence settings.
Persistent user preferences / custom instructions
provisionalverified ~2 months agoGemini supports persistent personalization through 'Memory' (Gemini learns from past chats — the successor to 'Saved info', which current help pages no longer use as a name) and 'Instructions for Gemini' (custom instructions that apply to every chat). Both are managed in the 'Personal Intelligence' section of Settings.
Plans and pricing
canonicalverified 26 days agoFour consumer tiers as of 2026-06-12: Free ($0), Google AI Plus ($4.99/mo, 2x Free usage — cut from $7.99 in early June 2026, with storage doubled from 200 GB to 400 GB), Google AI Pro ($19.99/mo, 4x Free usage), Google AI Ultra (starting at $99.99/mo for 5x Pro usage, with a $199.99/mo tier for 20x Pro usage).
Public share links
provisionalverified ~2 months agoAny Gemini chat can be shared as a public, link-only snapshot at a g.co/gemini/share/... URL. The shared page renders the full conversation as it existed at link-creation time, including Canvas docs, images, and generated videos. Recipients without a Google account can view; signed-in users (18+, non-Gem chats) can continue the chat in their own Gemini Apps.
Study notebooks: quiz-driven personalized lessons from uploaded materials
provisionalverified 25 days agoStudy notebooks are a dedicated notebook type (distinct from general context-grouping notebooks) that turns a user's own notes or course materials into an interactive, personalized study experience: diagnostic quizzes identify knowledge gaps, Gemini then builds bite-sized interactive lessons targeted at those gaps, and progress is tracked automatically as the learner completes activities.
Video generation (Gemini Omni / Veo)
provisionalverified ~2 months agoGemini generates video clips through the Gemini Omni multimodal model (announced at I/O 2026), which Google is rolling out to replace Veo as the create-and-edit-video model in the consumer Gemini app. Veo 3.1 / Veo 3.1 Fast / Veo 3.1 Lite remain available via the Gemini API and Google Flow as the prior-generation generators being phased out of the app.
Web search (Google Search grounding)
provisionalverified ~2 months agoGemini grounds its answers in real-time Google Search results when a query benefits from current information, surfacing inline citations and a 'Sources and related content' panel so users can verify claims. This is Gemini's everyday web-access capability, distinct from the multi-step Deep Research agent.