jnachi
Learning Hub
Growth & Problem Solving6 min readIntermediate

Trying Multimodal Features on Purpose (voice, image, and document inputs beyond plain text)

Expand beyond text prompts by integrating voice dictation, whiteboard screenshots, UI audits, and multi-page PDFs.

Works with:ChatGPT-4oClaude 3.5 SonnetGemini 1.5 ProCopilot

Key Takeaways

  • Speaking is ~3x faster than typing — use voice for unfiltered brain dumps
  • Screenshots communicate spatial, visual, and layout information that text cannot
  • Upload full PDFs to cross-reference across sections rather than copy-pasting excerpts
  • Multimodal prompts eliminate hours of manual transcription of whiteboards and diagrams
  • Error code screenshots beat typed descriptions for debugging — the model sees the exact syntax

The Diagnostic Context

When most people get stuck on a problem, they type long, exhaustive paragraphs into a chat window trying to describe a visual layout, a handwritten diagram, or a complex spreadsheet. This is the slowest possible way to communicate context. Modern multimodal models can see, hear, and parse visual relationships natively. Using image and voice inputs directly bypasses hours of manual transcription and explanation.

The Core Technique

Multimodality means using the right sensory medium for the information you possess:

1. Visual Synthesis (Screenshots & Diagrams)

  • Whiteboard to Jira: Take a phone snapshot of a messy whiteboard brainstorm from a conference room. Prompt: "Extract all sticky notes and grouped themes into a prioritized markdown backlog with suggested task owners."
  • UI/UX Critique: Take a screenshot of a live web page or Figma mockup. Prompt: "Analyze this landing page layout. Flag where visual hierarchy is broken, where text contrast fails accessibility standards, and where the primary call to action gets lost."
  • Error Code Diagnosis: Screenshot a terminal error, software stack trace, or messy Excel formula error rather than typing out the syntax manually.

2. Voice Interaction (Unfiltered Brain Dumps)

Speaking is roughly 3x faster than typing. Use real-time voice mode when walking or between meetings to ramble unstructured thoughts.

  • Voice Prompt: "I’m going to talk through our team's hiring priorities for 3 minutes without filtering. Listen to everything I say, filter out the tangents, and return a clean 1-page job specification."

3. Dense Document Grounding

Instead of copying and pasting sections, upload full PDFs (technical manuals, vendor contracts, academic papers). Use the model to cross-reference between distinct sections:

  • Directive: "Compare section 4.2's termination clause with the indemnity obligations in Appendix B."
5-Minute Activation Challenge

Try This Right Now

Take a photo or screenshot of something visual right now—a slide from a presentation, a messy desk whiteboard, a handwritten note, or an app screen. Upload it to your AI tool and ask: "Explain what this is showing, identify the 2 most important elements, and suggest one way to improve its clarity."

Tip: Knowledge only becomes capability once you run the prompt yourself.

Comprehension Check

Test Your Instincts (3 Questions)

1

When is uploading a screenshot superior to typing a text description into an AI prompt?

2

What is the primary operational benefit of using voice input modes with AI assistants?

3

Which task demonstrates effective use of document-based multimodal prompting?