How to build an AI agent with vision powered by Playwright

Suppose you want to build an agent that can handle this request:
@bot Redesign https://\[your-existing-landing-page\].com. Keep the content and brand. Use the attached image as a reference for the Features and Use Cases sections.
If you use a static scraper you'll only get HTML and CSS sent by the server. It may miss content added by JavaScript. It also cannot see the final spacing or how the layout changes on a phone.
With Playwright you can add these 3 useful inputs:
  • content added after the page loads
  • exact styles and section sizes from the rendered page
  • desktop and mobile screenshots

Contents

  1. Scrape a static page
  1. Render the page with Playwright
  1. Turn the design into structured data
  1. Send clear inputs to the model
  1. Capture interactive states and handle browser failures
  1. Compare static-only and Playwright-assisted redesigns

1. Scrape a static page

Run a static scraper before you open a browser. Accept only public HTTP and HTTPS URLs. Check every redirect and block private, local, and cloud metadata addresses. Add limits for response size, content type, and time.
The extraction uses fixed rules and no model call. Cheerio reads tags such as title, h1-h3, img. The structure parser scans section and article blocks, then groups headings, paragraphs, repeated cards, CTA links, and images
A shortened source snapshot looks like this:
JSON

Loading editor...

{
  "finalUrl": "<https://example.com>",
  "title": "AskDoc | AI assistants from company knowledge",
  "headings": [
    {
      "level": 1,
      "text": "Build AI assistants from your company knowledge"
    },
    ...
  ],
  "importantLinks": [
    {
      "text": "Create your first bot",
      "url": "..."
    },
    ...
  ],
  "sections": [
    {
      "kind": "hero",
      "order": 0,
      "heading": "Build AI assistants from your company knowledge",
      "supportingCopy": "Upload docs, guides, policies, and playbooks, then publish helpful assistants for customers, teams, and workflows."
    },
    {
      "kind": "use-cases",
      "order": 2,
      "heading": "Put your knowledge where the questions happen",
      "items": [
        { "title": "Customer support" },
        { "title": "Website conversion" },
        { "title": "Internal knowledge" }
      ]
    },
    ...
  ],
  "branding": {
    "colorScheme": "light",
    "roles": {
      "primary": { "value": "#1d867a" },
      "accent": { "value": "#f6a523" },
      "background": { "value": "#f9fafb" }
    }
  }
}
The scraper collects:
  • title, headings, and body text;
  • section order;
  • links and CTA targets;
  • product facts, prices, and other numbers;
  • logo, colors, and fonts.
Keep this snapshot as the main source for copy and facts. Text from a screenshot can be incomplete, and the model may read it wrong. The snapshot also gives you a fallback when browser capture fails.

2. Render the page with Playwright

Open the accepted URL in 2 fresh Playwright contexts:
  • desktop at 1440 × 1000;
  • mobile at 390 × 844.
This shortened helper opens a fresh context, loads the accepted URL, reads the rendered page, and takes a full-page PNG.
TSX

Loading editor...

import { chromium, type Browser } from "playwright";

async function captureViewport(
  browser: Browser,
  url: string,
  viewport: { width: number; height: number },
) {
  const context = await browser.newContext({
    viewport,
    acceptDownloads: false,
    permissions: [],
    serviceWorkers: "block",
  });

  try {
    const page = await context.newPage();
    await page.goto(url, { waitUntil: "domcontentloaded" });
    await page.evaluate(async () => {
      await document.fonts.ready;
    });

    return {
      sections: await readSections(page),
      screenshot: await page.screenshot({ fullPage: true }),
    };
  } finally {
    await context.close();
  }
}

async function capturePage(url: string) {
  const browser = await chromium.launch({ headless: true });

  try {
    return {
      desktop: await captureViewport(
        browser,
        url,
        { width: 1440, height: 1000 },
      ),
      mobile: await captureViewport(
        browser,
        url,
        { width: 390, height: 844 },
      ),
    };
  } finally {
    await browser.close();
  }
}

const captures = await capturePage(acceptedUrl);
Wait for the DOM, fonts, useful images, and a short period with no DOM changes. Avoid waiting for full network idle. Analytics and live requests can keep a page busy forever.
Capture one PNG for each viewport. The 2 screenshots show the model how the design changes across screen sizes.

3. Turn the design into structured data

Screenshots show the page, but they do not give the model exact values for spacing, section order, or link targets. Extract those details from the rendered DOM into Design JSON.
readSections runs in the page through Playwright. It reads each section's position and a small set of computed styles. The project runs it at both viewport sizes, matches the sections, and adds the result to the static snapshot.
TSX

Loading editor...

import type { Page } from "playwright";

async function readSections(page: Page) {
  return await page.locator("main section").evaluateAll((sections) =>
    sections.slice(0, 14).map((section, order) => {
      const box = section.getBoundingClientRect();
      const style = getComputedStyle(section);

      return {
        order,
        heading:
          section.querySelector("h1, h2, h3")?.textContent?.trim() ?? null,
        boundingBox: {
          x: Math.round(box.x),
          y: Math.round(box.y + window.scrollY),
          width: Math.round(box.width),
          height: Math.round(box.height),
        },
        styles: {
          display: style.display,
          backgroundColor: style.backgroundColor,
          color: style.color,
          fontFamily: style.fontFamily,
        },
      };
    }),
  );
}
Our Design JSON keeps:
  • section roles and bounding boxes;
  • type styles and colors;
  • navigation and footer data;
  • buttons, links, images, and controls;
  • desktop and mobile layout differences.
A shortened pricing section from the AskDoc Design JSON looks like this:
JSON

Loading editor...

{
  "designVersion": 1,
  "captureMode": "rendered",
  "viewports": [
    {
      "name": "desktop",
      "width": 1440,
      "height": 1000,
      "documentHeight": 4347
    },
    {
      "name": "mobile",
      "width": 390,
      "height": 844,
      "documentHeight": 7126
    }
  ],
  "sections": [{
    "id": "source-section-5",
    "kind": "pricing",
    "order": 4,
    "boundingBox": {
      "x": 0,
      "y": 3429,
      "width": 1440,
      "height": 702
    },
    "surfaceRole": "light",
    "headings": [
      "Pricing that fits a launch path",
      "Free",
      "Pro",
      "Business"
    ],
    "components": ["pricing", "cta"],
    "responsiveVariants": [
      {
        "viewport": "desktop",
        "boundingBox": {
          "x": 0,
          "y": 3429,
          "width": 1440,
          "height": 702
        }
      },
      ...
    ],
    "representativeStyles": {
      "section": {
        "backgroundColor": "rgb(255, 255, 255)",
        "color": "rgb(23, 27, 38)"
      },
      "heading": {
        "fontFamily": "Inter, ui-sans-serif, system-ui, sans-serif",
        "fontSize": "30px",
        "fontWeight": "600",
        "lineHeight": "36px"
      }
    }
  }]
}
Keep the data small. The project stores no more than 14 sections and limits Design JSON to 36 KiB. It leaves out the full DOM, scripts, class names, and browser state.

4. Send clear inputs to the model

Send 4 kinds of input in one generation request:
  1. the user's redesign request;
  1. copy, facts, links, and sections from the static snapshot;
  1. the Design JSON in the same bounded source payload;
  1. the desktop and mobile screenshots.
Tell the model which input controls each part of the result. The static snapshot controls copy, facts, and links. The screenshots and Design JSON guide layout and style. The user's request controls what should change.
Ask the model to return one complete HTML document. Then check the HTML against your security rules and compare it with the source facts. If the first result fails, send the exact errors back for one repair attempt.

5. Capture interactive states and handle browser failures

Run Playwright after the static safety checks. Use a fresh browser context with no cookies, storage, permissions, downloads, or access to private networks.
If Playwright times out or returns a bad screenshot, continue with the static snapshot. After generation, clear the screenshot buffers. Store the bounded source snapshot, which contains the static fields, Design JSON, diagnostics, and metadata. Log safe metrics such as capture time and image size.
After the first capture, the app finds labeled tabs, accordions, disclosures, carousels, and “show more” controls. It clicks them in page order and keeps only useful changes that stay within the same section and URL. For each accepted state, it sends small Design JSON changes and a clipped screenshot to the model. It restores the starting state or reloads the approved URL before it continues.

Here's the example of page missing content from static scraper. Playwright helped to preserve content and structure.

Ship faster than
your competition

Focus on customers and sales while we handle product delivery. Hire a dedicated AI maker or a whole product team.

Launch new products
Fix your delivery
Hit fundraising milestones
[Analytics Chart]
Wireframe v0.1
User
Sarah Chen CEO, Founder
$127k
MRR
2,847
Customers
+23%
Growth
Sales Performance
Latest Deal
Customer
Enterprise plan closed
Acme Corp · $12k/mo
© 2026 Paralect, Inc 651 N Broad St, Suite 206, Middletown, 19709, Delaware, United States