Multimodal Attachments Analysis for Auto Assist | The place for Zendesk users to come together and share
Skip to main content
Feedback submitted

Multimodal Attachments Analysis for Auto Assist

Related products:AI
  • June 29, 2026
  • 0 replies
  • 28 views

Paul30

Please give a quick overview of your product feature request or feedback and note who in your organization is affected by this issue [ex. agents, admins, customers, etc.] (2-3 sentences)

We are requesting that Zendesk Agent Copilot Auto Assist be updated with multimodal vision capabilities so it can read and analyze ticket attachments like photos, screenshots, and videos. This limitation directly impacts our front-line support agents who rely on customer-uploaded media to troubleshoot, as well as our admins who design Auto Assist Procedures. Adding vision support will ensure Auto Assist maintains context when critical customer data is sent as an image rather than text.

 

What problem do you see this solving? (1-2 sentences)

This feature solves the context gap that occurs when Auto Assist falls "one step behind" because it cannot process customer-submitted images, such as product serial number labels or error screenshots. By allowing Auto Assist to ingest visual attachments, the AI can seamlessly complete procedural workflows that depend on visual verification without breaking the suggested reply flow.

 

When was the last time you were affected by this lack of functionality, or specific tool? What happened? How often does this problem occur and how does this impact your business? (3-4 sentences)

This issue affected us today and continues to impact approximately half (50%) to two-thirds (66%) of our daily incoming support tickets where serial number photos are requested. When a customer sends a photo of their product label, Auto Assist misses the visual details entirely, proceeds with the wrong text-only steps, and throws off the procedural flow. This constant disruption forces our agents to frequently dismiss or completely disable Auto Assist, which dramatically slows down handling times and reduces our overall return on investment in the Agent Copilot suite.

 

Are you currently using a workaround to solve this problem? (If yes, please explain) (1-2 sentences)

Yes, our current workaround requires agents to manually download the customer's attachment, open the image to visually extract the product details, and manually type a reply based on that information back into the ticket or custom fields. Because Auto Assist is blind to the attachment and falls out of sync, agents have to dismiss the AI and take over the interaction entirely.

 

What would be your ideal solution to this problem? How would it work or function? (1-2 sentences)

Our ideal solution is a multimodal Auto Assist where the underlying model can detect and process visual media attachments via instructions built directly into procedures (e.g., "Analyze the attached photo to extract the serial number and model"). Auto Assist would then automatically extract the information, populate the corresponding custom fields, and suggest the correct next troubleshooting step without breaking stride.