Before filing
Closest existing issue
none found
Is this new, or an improvement?
New capability — Berd can't do this at all today
The problem, in your terms
I want Berd agents to interact with a computer and browser as part of their normal work—for example, opening websites, navigating pages, entering text, clicking controls, inspecting rendered pages and completing browser-based tasks.
Buzz.xyz already provides a computer-control/browser automation capability. Because Buzz.xyz and Berd.xyz are both Block projects, I expected the same capability to be available in Berd. At present, I cannot use that functionality directly from a Berd agent.
This limits Berd agents to files, repositories and other configured tools, while preventing them from completing tasks that require visual browser or computer interaction.
What you do today
I would need to use Buzz separately or install and configure a third-party local browser/computer-control MCP server manually. This creates another silo and prevents Berd agents from using the same computer-control capability within their existing projects and chats.
What you'd like to see
Buzz’s repository includes buzz-dev-mcp, and Buzz exposes agent tools for interacting with the workspace. Please consider making the relevant computer/browser-control capability available in Berd through an equivalent native MCP integration, using the same underlying implementation where practical but adapting execution and permissions for local macOS use.
Ideally, Berd agents should be able to:
- Open and navigate websites.
- Click, type, scroll and select elements.
- Inspect page content and rendered state.
- Upload and download files where permitted.
- Use browser sessions associated with a project or chat.
- Receive screenshots or page-state information.
- Ask for confirmation before sensitive actions.
- Work with local browser automation on macOS.
- Configure permissions per agent and project.
Berd should support screen recording and workflow capture so a user can demonstrate a task once, have Berd convert it into a reusable agent skill, and then allow authorised agents to execute that skill. This should include the relevant screen, mouse, keyboard, application and browser context, with clear privacy controls, review/editing before skill creation, and confirmation before replaying sensitive actions.
This capability should be comparable to OpenAI Codex Record & Replay and/or Hermes /learn skill while integrating with Berd’s existing agents, projects, skills and MCP tools.
The implementation could use the same underlying technology as Buzz where practical, but the user experience should be native to Berd.
Why this belongs in Berd itself
This is a core capability for desktop agents that are expected to work across projects and tools. A third-party MCP server could provide browser automation, but users would have to discover, install, secure and configure it separately.
A built-in integration would provide:
- Consistent installation.
- Secure credential and permission handling.
- Native project and agent access controls.
- Better macOS integration.
- A consistent experience across Block’s agent products.
Because Buzz already provides related functionality, sharing or adapting the underlying implementation could also reduce duplication across Block projects.
Non-goals
- This request is not asking Berd to become a full remote Buzz relay.
- It is not asking to merge Berd and Buzz into one product.
- It is not asking for unrestricted autonomous computer access.
- It is not asking for surveillance.
- It is not asking for automatic actions without user permissions.
- It is not asking for every browser automation framework to be bundled.
The screen recording is not a request for continuous background surveillance or automatic recording. Recording should be explicitly started by the user, visibly active, stoppable at any time, and subject to local privacy controls.
This would make the feature request stronger because it asks not only for computer control, but also for turning demonstrated workflows into reusable Berd skills.
Alternatives you considered
No response
Mockups, prior art, or other context
Buzz.xyz
Buzz GitHub repository
Buzz architecture documentation
Model Context Protocol tools specification
Berd GitHub repository
Before filing
Closest existing issue
none found
Is this new, or an improvement?
New capability — Berd can't do this at all today
The problem, in your terms
I want Berd agents to interact with a computer and browser as part of their normal work—for example, opening websites, navigating pages, entering text, clicking controls, inspecting rendered pages and completing browser-based tasks.
Buzz.xyz already provides a computer-control/browser automation capability. Because Buzz.xyz and Berd.xyz are both Block projects, I expected the same capability to be available in Berd. At present, I cannot use that functionality directly from a Berd agent.
This limits Berd agents to files, repositories and other configured tools, while preventing them from completing tasks that require visual browser or computer interaction.
What you do today
I would need to use Buzz separately or install and configure a third-party local browser/computer-control MCP server manually. This creates another silo and prevents Berd agents from using the same computer-control capability within their existing projects and chats.
What you'd like to see
Buzz’s repository includes buzz-dev-mcp, and Buzz exposes agent tools for interacting with the workspace. Please consider making the relevant computer/browser-control capability available in Berd through an equivalent native MCP integration, using the same underlying implementation where practical but adapting execution and permissions for local macOS use.
Ideally, Berd agents should be able to:
Berd should support screen recording and workflow capture so a user can demonstrate a task once, have Berd convert it into a reusable agent skill, and then allow authorised agents to execute that skill. This should include the relevant screen, mouse, keyboard, application and browser context, with clear privacy controls, review/editing before skill creation, and confirmation before replaying sensitive actions.
This capability should be comparable to OpenAI Codex Record & Replay and/or Hermes /learn skill while integrating with Berd’s existing agents, projects, skills and MCP tools.
The implementation could use the same underlying technology as Buzz where practical, but the user experience should be native to Berd.
Why this belongs in Berd itself
This is a core capability for desktop agents that are expected to work across projects and tools. A third-party MCP server could provide browser automation, but users would have to discover, install, secure and configure it separately.
A built-in integration would provide:
Because Buzz already provides related functionality, sharing or adapting the underlying implementation could also reduce duplication across Block projects.
Non-goals
The screen recording is not a request for continuous background surveillance or automatic recording. Recording should be explicitly started by the user, visibly active, stoppable at any time, and subject to local privacy controls.
This would make the feature request stronger because it asks not only for computer control, but also for turning demonstrated workflows into reusable Berd skills.
Alternatives you considered
No response
Mockups, prior art, or other context
Buzz.xyz
Buzz GitHub repository
Buzz architecture documentation
Model Context Protocol tools specification
Berd GitHub repository