Kaizen Teams

Dropdown

Table of Contents

Time to read

·

12

Published on

·

January 22, 2026

Last updated on

·

August 27, 2026

Esteban Silva, Full-Stack Developer at Kaizen Softworks

Esteban Silva

No in-between guy

Full-Stack Developer

AI

AI

Technology

Technology

Building a computer vision POC in hours with Vercel AI SDK and Figma MCP

Published on

·

August 27, 2026

Last updated on

·

August 27, 2026

Time to read

·

12

Esteban Silva, Full-Stack Developer at Kaizen Softworks

Esteban Silva

Full-Stack Developer

"What if we could build an app in days instead of weeks, using AI not just as a feature, but as a full-stack development partner?"

That's the question that sparked Super Ultra Kiosco, a proof of concept (POC) that digitized our company's beloved manual snack corner using the latest AI development tools.

We sat down to document exactly how we used tools like Figma MCP (Model Context Protocol) and Vercel AI SDK to build a working product detection system.

Let's start with context: What exactly is "Super Ultra Kiosco"?

In our office, "El Kiosco" is a small corner with snacks, drinks, and other goodies. The system was manual: grab an item, find the shared spreadsheet, and manually log your debt. It was error-prone, tedious, and ripe for modernization.

The Solution: A mobile-first web app where you simply point your phone at the snacks. AI identifies the products via camera, calculates the total, and confirms your order with one tap. No spreadsheets, no typing.

Mobile interface of Kiosko AI, a computer vision POC built with Vercel AI SDK and Figma MCP, demonstrating real-time product recognition.

The tech stack behind this AI-native app

To achieve this, we moved beyond standard web development into an AI-native workflow. Here is the core technology stack:

  • Framework: Next.js with TypeScript.
  • AI Vision & Logic: Vercel AI SDK (powered by Google Gemini Vision) for the product detection feature itself.
  • Design-to-Code: Figma MCP for translating designs into code.

This combination allowed us to treat the AI as a collaborator that could "see" our designs and "understand" our business logic.

How does product detection with Vercel AI SDK work?

The "magic" of the app is its ability to recognize a bag of chips or a soda can instantly. For this, we used the Vercel AI SDK coupled with Google’s Gemini Vision model.

Traditional AI integrations often struggle with unstructured text responses. We solved this using the SDK's generateObject function. This allows you to define a schema (using Zod) to force the AI to return structured, typed JSON data instead of a conversational string.

Here is the actual implementation code:

JavaScript

const result = await generateObject({

  model: google('gemini-2.5-flash'),

  messages: [

    {

      role: 'user',

      content: [

        { type: 'text', text: prompt },

        { type: 'image', image }

      ]

    }

  ],

  schema: AnalysisResultSchema,

  temperature: 0.1

});

The Trade-off: Using the AI SDK provides incredible flexibility to switch model providers (e.g., swapping Gemini for GPT-4o) without rewriting code. However, if you need bleeding-edge features specific to a provider—like accessing a brand new beta model or a very specific configuration—you might find it limited. You gain high availability of providers, but you might have a smaller pool of models available within each provider compared to using their native SDKs.

What about Figma MCP? How did that fit into the workflow?

Figma MCP (Model Context Protocol) creates a bridge between your Figma designs and your AI coding assistant (like Windsurf or Cursor).

In technical words, Figma MCP is a standard that connects AI assistants directly to external data sources. The Figma MCP server exposes your actual design files—component properties, layout tokens, and visual hierarchy—as a resource the AI can "read."

​​Instead of us manually coding CSS from mockups, the workflow looked like this:

  1. We designed the UI in Figma.
  2. Our IDE (via MCP) "connected" to the design file.
  3. We prompted the AI: "Build this product card component based on the 'Mobile Card' frame in Figma."
  4. The AI generated React components that matched our design specs (almost) perfectly.
Order confirmed UI screen for Kiosko AI, signaling a successful computer vision transaction built with Vercel AI SDK and Figma MCP.

Key Learning: While Figma MCP helped us move fast, it wasn't magic. We found that we still had to iterate a few times to match the designs 100%. The AI captures the general structure well, but it can fail on specifics like exact paddings, button states, or color nuances. It gets you very close, but the final polish to ensure the implementation matches the design perfectly still requires a developer's eye.

What was the most surprising part of building with these AI tools?

Building Super Ultra Kiosco validated four core hypotheses about the state of AI development:

  • Speed: We moved from concept to working prototype in just 6 hours, a fraction of the time it would normally take. The structured output from Vercel’s SDK removed the need for complex parsing logic.
  • Context is everything: When using AI assistants for development, providing good context (whether through MCP or clear instructions) dramatically improves results. With AI SDK giving us structured outputs and Figma MCP ensuring design fidelity, we spent less time on boilerplate and debugging, and more time on the actual product experience.
  • Prompt engineering is still an art: Even with structured outputs, crafting the right prompt for product detection took iteration.
  • Image quality matters: The AI detection works best with good lighting and clear photos. We added image compression to balance quality vs. API costs.

Q&A: Common questions on AI-assisted development

Is Figma MCP ready for production apps? It is excellent for rapid prototyping and setting up component libraries. For complex, custom animations or highly specific accessibility requirements, human oversight is still mandatory.

Why use Vercel AI SDK instead of the official Google SDK? For a POC or a multi-model application, Vercel AI SDK offers a unified API that saves significant time. If your app relies 100% on unique, deep features of a single model (like Gemini's 1M context window specific caching), the native SDK might be better.

What is the cost implication of using Vision models for this? Vision models are more expensive than text models. We implemented client-side image compression before sending requests to the API to balance performance and cost.

Building something similar? 

We'd love to hear about your experience with AI development tools.

LET’S TALK

"What if we could build an app in days instead of weeks, using AI not just as a feature, but as a full-stack development partner?"

That's the question that sparked Super Ultra Kiosco, a proof of concept (POC) that digitized our company's beloved manual snack corner using the latest AI development tools.

We sat down to document exactly how we used tools like Figma MCP (Model Context Protocol) and Vercel AI SDK to build a working product detection system.

Let's start with context: What exactly is "Super Ultra Kiosco"?

In our office, "El Kiosco" is a small corner with snacks, drinks, and other goodies. The system was manual: grab an item, find the shared spreadsheet, and manually log your debt. It was error-prone, tedious, and ripe for modernization.

The Solution: A mobile-first web app where you simply point your phone at the snacks. AI identifies the products via camera, calculates the total, and confirms your order with one tap. No spreadsheets, no typing.

Mobile interface of Kiosko AI, a computer vision POC built with Vercel AI SDK and Figma MCP, demonstrating real-time product recognition.

The tech stack behind this AI-native app

To achieve this, we moved beyond standard web development into an AI-native workflow. Here is the core technology stack:

  • Framework: Next.js with TypeScript.
  • AI Vision & Logic: Vercel AI SDK (powered by Google Gemini Vision) for the product detection feature itself.
  • Design-to-Code: Figma MCP for translating designs into code.

This combination allowed us to treat the AI as a collaborator that could "see" our designs and "understand" our business logic.

How does product detection with Vercel AI SDK work?

The "magic" of the app is its ability to recognize a bag of chips or a soda can instantly. For this, we used the Vercel AI SDK coupled with Google’s Gemini Vision model.

Traditional AI integrations often struggle with unstructured text responses. We solved this using the SDK's generateObject function. This allows you to define a schema (using Zod) to force the AI to return structured, typed JSON data instead of a conversational string.

Here is the actual implementation code:

JavaScript

const result = await generateObject({

  model: google('gemini-2.5-flash'),

  messages: [

    {

      role: 'user',

      content: [

        { type: 'text', text: prompt },

        { type: 'image', image }

      ]

    }

  ],

  schema: AnalysisResultSchema,

  temperature: 0.1

});

The Trade-off: Using the AI SDK provides incredible flexibility to switch model providers (e.g., swapping Gemini for GPT-4o) without rewriting code. However, if you need bleeding-edge features specific to a provider—like accessing a brand new beta model or a very specific configuration—you might find it limited. You gain high availability of providers, but you might have a smaller pool of models available within each provider compared to using their native SDKs.

What about Figma MCP? How did that fit into the workflow?

Figma MCP (Model Context Protocol) creates a bridge between your Figma designs and your AI coding assistant (like Windsurf or Cursor).

In technical words, Figma MCP is a standard that connects AI assistants directly to external data sources. The Figma MCP server exposes your actual design files—component properties, layout tokens, and visual hierarchy—as a resource the AI can "read."

​​Instead of us manually coding CSS from mockups, the workflow looked like this:

  1. We designed the UI in Figma.
  2. Our IDE (via MCP) "connected" to the design file.
  3. We prompted the AI: "Build this product card component based on the 'Mobile Card' frame in Figma."
  4. The AI generated React components that matched our design specs (almost) perfectly.
Order confirmed UI screen for Kiosko AI, signaling a successful computer vision transaction built with Vercel AI SDK and Figma MCP.

Key Learning: While Figma MCP helped us move fast, it wasn't magic. We found that we still had to iterate a few times to match the designs 100%. The AI captures the general structure well, but it can fail on specifics like exact paddings, button states, or color nuances. It gets you very close, but the final polish to ensure the implementation matches the design perfectly still requires a developer's eye.

What was the most surprising part of building with these AI tools?

Building Super Ultra Kiosco validated four core hypotheses about the state of AI development:

  • Speed: We moved from concept to working prototype in just 6 hours, a fraction of the time it would normally take. The structured output from Vercel’s SDK removed the need for complex parsing logic.
  • Context is everything: When using AI assistants for development, providing good context (whether through MCP or clear instructions) dramatically improves results. With AI SDK giving us structured outputs and Figma MCP ensuring design fidelity, we spent less time on boilerplate and debugging, and more time on the actual product experience.
  • Prompt engineering is still an art: Even with structured outputs, crafting the right prompt for product detection took iteration.
  • Image quality matters: The AI detection works best with good lighting and clear photos. We added image compression to balance quality vs. API costs.

Q&A: Common questions on AI-assisted development

Is Figma MCP ready for production apps? It is excellent for rapid prototyping and setting up component libraries. For complex, custom animations or highly specific accessibility requirements, human oversight is still mandatory.

Why use Vercel AI SDK instead of the official Google SDK? For a POC or a multi-model application, Vercel AI SDK offers a unified API that saves significant time. If your app relies 100% on unique, deep features of a single model (like Gemini's 1M context window specific caching), the native SDK might be better.

What is the cost implication of using Vision models for this? Vision models are more expensive than text models. We implemented client-side image compression before sending requests to the API to balance performance and cost.

Building something similar? 

We'd love to hear about your experience with AI development tools.

LET’S TALK

Related Articles

View all articles

·

Aug 28, 2026

About Catalyst 26: Partnerships & Ecosystem Conference

Everything to know about Catalyst 26: dates, price, who attends, both keynote recaps, and when the next Catalyst event is.

12 read time

Read more

Catalyst 26 was Partnership Leaders' fifth annual conference for partnership, ecosystem, and go-to-market professionals. It took place August 25 and 26, 2026, at the Marriott Hotel at the Brooklyn Bridge in New York, with more than 1,000 attendees and 70-plus speakers from companies including Anthropic, OpenAI, Google, Microsoft, IBM, BCG, and Siemens.

Dates August 25–26, 2026
Location Marriott Hotel at the Brooklyn Bridge, Brooklyn, NY
Edition 5th annual
Attendees 1,000+ partnership, ecosystem, and GTM professionals
Speakers 70+, including people from Anthropic, OpenAI, Google, Microsoft, IBM, BCG, and Siemens.
Price $849 early bird, rising to $999, then $1,999

Who Catalyst events are for

Catalyst brought together people building and running partner programs across SaaS, AI, consulting, systems integration, agencies, and major cloud platforms.

Attendees included executives leading partnership organizations, and people working directly in alliances, partner sales, marketing, operations, strategy, and enablement.

What Catalyst 26 is like

You can look at the agenda before a conference and have a pretty good idea of what you'll find. Being there is different.

This year's theme was "Navigating Frontier Ecosystems". Anthropic's Head of Partnerships and one of OpenAI's partner program leads appeared on the same agenda as people from Oracle, Siemens, IBM, and BCG, companies that have run formal partner programs for two decades.

That mix was one of the most interesting parts of the conference. Newer AI companies were discussing partner tiers, co-selling, and joint delivery alongside companies where those models have been part of their business for years.

What Catalyst 26 covered

Catalyst 26 split its sessions into eight pillars:

  • Advancing Organizational Maturity: turning partnerships into something measured and repeatable instead of one founder doing favors for another.
  • Become a Strategic Partner: getting partnerships involved when product and business decisions are made, not told about them afterward.
  • Frontier Partner Experience: adapting partner programs as AI changes how companies build and integrate products.
  • Path to CPO: career sessions for people aiming to lead partnerships at the executive level.
  • Co-Build: two companies building something together.
  • Co-Market: two companies running a campaign together.
  • Co-Sell: two sales teams working the same deal.
  • Co-Serve: two companies delivering the same engagement to a client.

Catalyst 26 sessions

Day 1 Keynote

The Day 1 keynote brought together Partnership Leaders’ CEO Asher Mathew, Tribe AI’s Co-founder & CEO Jaclyn Rice Nelson, Anthropic’s Head of Partnerships Phil Samenuk, and Boomi’s Chairman & CEO Steve Lucas.

Their discussion focused on how companies are relying on partners to build, sell, and deliver products across AI, cloud, and enterprise software. A few points stood out:

  • More companies have dedicated partner teams now, which means a generic, one-size-fits-all partner program doesn't cut it anymore. Partners show up when the program fits how they work.
  • New AI products and cloud services are shipping so fast that a partner program can't just get set once and left alone. Incentives, support, and how you work together need regular updates.
  • Partnerships also came up as a way to access data a company couldn’t reach on its own, whether that meant getting access to it, combining it, or putting it to use.
  • AI doesn't change the basics of a good partnership. Account planning, clear ownership, and relationships built over time still matter most.

Day 2 Keynote

The Day 2 keynote featured Ramp’s Lead Economist Ara Kharazian, Eliza’s Founder Brian Benedict, Siemens’ EVP Global Partner Ecosystem Dion Smith, and Oracle’s SVP, Partner Sales & Operations Strategy Leah Yomtovian.

A few points stood out:

  • The spending data told a slower story than expected: AI adoption is mostly going toward productivity gains and task automation, not some overnight shift.
  • Siemens is in the middle of folding more than 68,000 partners and roughly 200 separate programs into a single global one, mainly to make it easier to coordinate across IT and operational technology.
  • Oracle's approach is a running "listening tour": every partner gets the same baseline benefits, then incentives and credits get layered based on the type of partner and how they work with Oracle.
  • There was also talk of a newer kind of service team: bring in engineers, turn AI requirements into working products, and reuse delivery methods that already work instead of starting from scratch each time.

Next Catalyst events

The date and location of Catalyst 27 hasn’t been announced yet. In the meantime, you can check out the half-day Catalyst Summits in different cities:

  • October 20, 2026 - Seattle
  • October 27, 2026 - Chicago
  • October 2026 - Los Angeles
  • December 2026 - Singapore

Check Partnership Leaders’ events page for updates.

·

Aug 26, 2026

Why adding people doesn't always fix a struggling team

Learn when a software team should hire, wait, reorganize, or build skills internally, and how to tell which option will actually help.

12 read time

Read more

When a client asks to hire someone new, a common reaction is to open a search. There's more work, more pressure, and new features to build. It seems like the obvious thing to do.

But in our experience working with software development teams, the problem often isn't a lack of people. The problem is knowledge concentrated in too few people, unclear team roles, slow onboarding, or temporary demand.

The question worth asking isn't who can fill the position, but what would help the team work better. That points to one of three answers: hire, don't hire, or build the capability from within. Figuring out which one applies, and why, is the real work before opening a search.

What you should ask before assuming you need someone new

Hiring works when three conditions are met: the need will last, no one on the team has the capacity to take it on, and the team can onboard someone well. That last condition is easy to overlook. A team can have a real, lasting gap and still not be ready to bring someone in if no one has the time to guide them.

The risk comes from jumping straight from "there's more work" to "we need someone" without checking what's causing the pressure. It's easy to turn a request into a list of requirements (X years of experience, a specific technology, advanced English) and start the search. The real cause is often something else: a project that grew too fast, a tech lead with no time to onboard new hires, processes that stopped scaling, or a team that lost key people and needs to recover knowledge before adding headcount.

That's why, before thinking about who could fill the role, we ask these questions:

  • What outcome is the client trying to achieve?
  • What's happening on that team today?
  • What specific problem is this hire meant to solve?
  • Does adding a person solve that problem?
  • Is there someone on the team who could take this on?
  • Are there other, less obvious alternatives?

When the answers confirm the need will last, the current team can't cover it, and the team has the capacity to onboard someone, hiring is the right call: opening the search fills a gap the team can't close internally.

Does the problem need someone new to fix it?

Not hiring is the right call when the problem behind the request is temporary, or when it will resolve before the new hire finishes onboarding. Recommending against a hire may sound unusual for a company that offers staff augmentation, but our job as a strategic partner is not to maximize every opportunity but to recommend the best decision for the client. Depending on what's actually going on, the fix can look like:

  • An internal rotation: moving someone with spare capacity into the gap.
  • Reorganizing responsibilities across the team instead of adding a seat.
  • Hiring a different profile than the one originally requested.
  • Combining two roles into one instead of opening two searches.
  • Waiting a few weeks, when the project context is about to change on its own.

Is a temporary increase in workload a good reason to hire?

This happened on a project with a long onboarding period. The initial request seemed clear: hire a mid-level developer. There was work and budget available. But when we spoke with the team, we found that the workload increased because one team member had been temporarily reassigned to another sub-team. Before moving forward, we considered what would happen when that person came back.

The client's system was complex: any new hire needed several months to understand the business, the architecture, and the platform before they could contribute independently.

The problem justifying the hire was going to disappear, but the new hire wouldn't. By the time that person had enough context, the need that started the search would no longer exist.

We recommended against moving forward, even though there was budget to add someone. The client avoided an unnecessary hire and months of onboarding for a problem that was already resolving itself. Sometimes the best answer is to wait a few weeks; other times, it's reorganizing the team or developing internal talent.

How can you build team capability without hiring?

Build capability internally when the team already has product context but lacks a specific skill. Developing that skill internally can be faster than waiting for someone new to reach the same level of context.

More people doesn't always mean more capacity. Onboarding a new hire takes time from the people already on the team: explaining the business and the architecture, reviewing their work, and building trust. That's why, during the first few weeks, a team can become less productive while it onboards someone new. Complex projects can include years of technical decisions and undocumented knowledge. New hires still need time to learn that context.

Should you hire a specialist or train someone on your team?

A client needed a senior SQL Server specialist. That niche skill set made the role difficult and expensive to fill. We started the search and interviewed candidates, but the deeper issue became clear quickly: the real challenge on the project wasn't SQL Server. It was understanding a product shaped by years of evolution, multiple applications, and complex business logic.

The right person to develop that expertise was already on the team. Instead of hiring someone with deep SQL Server expertise, the client supported that team member in building the SQL Server skills the project needed. That person had business knowledge, motivation, and a much shorter learning curve than an external hire would have had. An outside specialist provided targeted support when needed.

The team gained SQL Server expertise without losing months waiting for a new hire to learn the product first. The person who took on SQL Server gained a valuable new skill without stepping away from the other work they were doing on the project.

A team's capacity depends on how its people complement each other, what knowledge they share, and what autonomy they've developed, not just on headcount. A team of ten people who are aligned, with shared context and autonomy, can generate more value than a team of fifteen where much of the time goes into onboarding new hires.

Should you hire, wait, or develop the skill internally?

Scenario Signal What to do
Hire The need will last, no one on the team can cover it, and the team can onboard someone well Open the search for a clearly defined role
Don't hire The problem is temporary or resolves before onboarding finishes Wait, reorganize the team, or cover the gap another way
Build internal capability Missing specific expertise, not people; someone already has the business context Develop the skill internally, with targeted outside support if needed

What questions do we ask first?

  1. What specific problem are we trying to solve?
  2. Will the need still exist after the person has been hired and onboarded?
  3. Is there someone on the team who could cover it?
  4. Do we have the capacity to onboard someone well?
  5. Is the problem a lack of people, or is it caused by unclear roles, missing product knowledge, slow onboarding, or a temporary increase in workload?
  6. What impact will this hire have six months from now?
  7. If we couldn't hire today, what other option would we explore?
  8. What higher-priority work would someone on the team have to stop doing to cover this need?

Wait to open a search when the team can't define the problem, confirm the need will last, or support onboarding. Clarify those points first.

If you're weighing this decision with your own team, let's talk about whether to hire, reorganize, or develop someone already on the team.

llms.txt