One skill set, two agents: enabling Claude and MuleSoft Vibes across the development lifecycle

Rich Stephenson

11 mins read.

September 20, 2026

CONTENTS

The starting question

Claude vs. MuleSoft Vibes – which of these common AI tools is best for MuleSoft application generation? The original goal was a straight comparison, but this question was not the right one to ask.  

In a real project, you don’t standardize on a single AI tool for the whole lifecycle. You might document with one tool, generate with another and review with an opposing tool to the one that generated. If your coding standards live inside one tool’s config format, you’re maintaining multiple copies of the same rules.   

Or you may have a developer or client preference on which tool provides the best output, and this may not even be the AI tool that we started a comparison for.   

Add to this, the rate of recent change for all AI generative tools, it may be better to have a set of generic skills and processes that remain relevant with such advancement in the AI space. 

The real problem worth solving, then morphed into “Can one skill set be used to drive multiple AI tools or agents through generation, review, testing, and hardening without change?” 

Test scenario

To get a realistic testbed, we built a specification for AnyFreightFast, a fictional freight company, as a scenario with 5 applications spanning Experience, Process, and System API layers. Claude generated the scenario and a full specification.  

MuleSoft Vibes produced the first draft of a SKILL.md for building against it, grounded with MuleSoft’s Best Practices and Adaptiv’s own established MuleSoft’s Best Practices. Each tool then reviewed and revised the other’s output, and we built a test application from both,  

The generation quality was roughly comparable between tools with no clear winner. Both interpreted the specification well and had different solutions, but each required multiple reviews, which pointed to skills refinement. It also hinted at the process to come, that is, generation is not a one-step process, there must always be a review step. What the exercise also showed was a shared skill was viable, and the interesting engineering problem was building one set properly, not choosing which tool was better. 

What we had at that point: one SKILL.md of around 2,000 lines, covering code generation, MUnit generation, and CI/CD in a single file.

Refactor 1: strip the monolith

A 2,000-line skill file has two concrete costs: 

  • Context window bloat – the full file loads into the context window on every invocation of the skill, but only part of it may be needed for the current task.  This reduces the capacity of the context window for other tasks and slows down responses. 
  • Token cost – loading large skills files adds to the token cost, and therefore how much of your usage allowance a session consumes.  They can get consumed repeatedly for the session as all history is reprocessed, so avoid large skills files and skills you don’t need. 

The fix is to move reference material, examples, and reusable assets into subfolders, and keep only the operative instructions in SKILL.md itself. That took the main file from 2,000 lines to around 200 lines.

mulesoft-code-generation/
├── SKILL.md             # ~200 lines, operative instructions only
├── references/
│   └── coding-standards.md
├── examples/
│   └── sample-flow.xml
└── assets/
    └── template-pom.xml

The general guidance we’re working to now is to keep SKILL.md files under 500 lines. Past that, split it into sub folders and files rather than trim it further. 

Refactor 2: split by lifecycle stage

A single skill trying to cover generation, testing, and CI/CD gets loaded in full even when only one of those concerns applies, MUnit generation, for example, is only relevant at the end of a build-review cycle, not throughout it. We split the monolith into lifecycle-scoped skills: 

Skill Purpose
mulesoft-code-generation Generate a MuleSoft app in line with best practices and Adaptiv standards
mulesoft-code-review Run a standards-based code review
mulesoft-munit-generation Generate MUnit tests
mulesoft-munit-review Run a standards-based MUnit review
mulesoft-security-hardening Review against the OWASP Top 10 and related security concerns

Note also that the content of the skill is shared between skills and linked via file references. For example, the code review skill refers to the code generation skill’s content (it holds the standards, assets, references and examples), so these are not repeated in the review skill.

This means the set of skills need to exist locally as a group to work together. More of this below.

Splitting this way also enables sub-agent isolation – a code-review pass can run in an agent instance with zero context from the generation session. There is no shared memory of how the code was written, only the code itself and the review skill. This a deliberate choice to keep the review independent and reflect closer to how a human reviewer with no prior involvement would approach it.

An alternative approach is to have one tool review the other’s output. For example, have MuleSoft Vibes generate an application, and have Claude review the output.

What is important is to try and run the review in a way that it is not in the same context window as the one that produced the output, so the review is independent from generation.

Having separate skills for reviews (mulesoft-code-review and mulesoft-munit-review) allows for reviews on existing code as well as recently generated code, and to report only, not perform any generating actions.

Other skills were added to complement lifecycle requirements.

Skill Purpose
mulesoft-documentor Document an existing application
drawio AI-assisted draw.io diagram generation

You may notice a missing skill for deployment. We took this out of the original monolithic skill, as we have standard CI/CD processes in place (source control, including branching strategies, pull requests and code reviews) which we use as the final human-in-the-loop check on the code, and to manage deployments to the Anypoint Platform.

This deployment skill may come into action later as we become more confident with the generation skills output, or align to the review skills, ensuring no blocker or major issues before automating check-ins and deployments. However, at present we are more comfortable with the consistency and governance of existing CI/CD processes for deployments.

Skills Refinement

A final step was a team exercise in using the new skills via Claude and MuleSoft Vibes, with refinements such as logging and flow / sub-flow generation that were identified. Changes were peer reviewed via pull requests, incorporating new additions into the Skills repositories. 

We see this as an ongoing process, to improve the skills that we use for reviews and generation.

Why a shared skill set is worth the engineering effort

  • Standards consistency – output converges on the same conventions regardless of which agent generated it. 
  • Single maintenance surface – one skill to update, not multiple tool-specific variants that drift apart over time. 
  • Tool flexibility – clients have tool preferences; a portable skill set decouples your standards from any one vendor’s assistant. 
  • Vibes uses Claude under the hood – as well as its own MuleSoft specific LLM. The two tools can generate fairly similar output when grounded with the same skills and templates, and the out-of-the-box MuleSoft Vibes skills are disabled. 
  • Format portability – the Skills format (SKILL.md + folder structure) is being adopted by other LLMs/agents, so it’s a reasonably safe bet against future tool changes, not a single-vendor lock in. 

Skills as versioned assets

Skills should live in a Git repository – we use Azure Dev Ops, but any is fine. This is just for good source control hygiene. We use a single repository containing all MuleSoft related skills, with a main branch, and feature or bugfix branches to add or upgrade skills. The workflow is: 

  1. Clone the repo locally and move skills folders to the appropriate location. 
  2. Use the required skills in generating and reviewing applications and MUnits. 
  3. If you find a fix or improvement when reviewing generated applications,  
  4. Branch and apply changes and raise a PR with the change for review. 
  5. Get it reviewed and merged and notify of an update. 
  6. Don’t just patch your local copy. 

This matters because local, un-committed fixes are invisible to everyone else on the next clone and so the entire point of a shared skill set collapses, if improvements don’t flow back.

Scope: global vs. project 

We settled on global skill installation over per-project. Reasoning: 

  • No need to check skill folders into every project repo. 
  • One canonical copy per developer machine, rather than multiple slightly diverged copies across projects. 

Mechanically, this differs by tool. Claude treats a folder drop into its global skills directory as a straightforward, no-config addition. Vibes has a different folder structure and requires going through its Skills panel where skills can be enabled or disabled.

Enabling a skill in MuleSoft Vibes 

  1. Clone the skill from ADO into a local folder. 
  2. In VS Code, open the MuleSoft Vibes extension → click the scales icon at the bottom of the panel → select Skills (not Rules – deprecated; not Workflows). 
  3. MuleSoft-provided skills are listed first and can be toggled on/off independently. 
  4. Locate the global skills folder: 
  5. Windows: C:\Users\{name}\AppData\Roaming\Code\User\globalStorage\salesforce.mule-dx-vscode\Skills 
  6. Mac: ~/Library/Application Support/Code/User/globalStorage/salesforce.mule-dx-vscode/Skills 
  7. Copy the entire skill folder into that directory – SKILL.md plus every subfolder (references/, examples/, assets/). 
  8. Back in Vibes, confirm the skill now appears under Global Skills. If it’s missing, check SKILL.md has a name and description defined at the top – Vibes won’t register a skill without both. 
  9. Verify by asking the agent directly: “What skills do you see?” 
  10. Note that within MuleSoft Vibes, there are several pre-configured skills, that may attribute to differing results.  All skills within Vibes have an on/off toggle if you wish to ignore any in code generation. 

Enabling a skill in Claude 

  1. Clone the skill from ADO into a local folder. 
  2. Copy the folder into C:\Users\{name}\.claude\skills 
  3. Windows: C:\Users\{name}\.claude\skills 
  4. Verify by asking the agent directly: “What skills do you see?” 
  5. If the new skill is not picked up, restart the session – Claude Desktop, Claude Code, or terminal. 

AGENTS.md: scope it correctly 

AGENTS.md is a markdown file that gives the AI tool the context and instructions for a specific project. It is not a skill and shouldn’t absorb skill-level content even though nothing technically stops you from putting it there.  

Its scope is narrow and project-specific, for example, project/application name, deployment target, and specific skills overrides, such as a runtime or connector version. 

The AGENTS.md file lives in the project root (e.g., {application}/AGENTS.md). 

Reference the specification or component that provides build detail, we should not just rely on a prompt to build the application. 

AGENTS.md is becoming an open convention, similar to the acceptance of SKILLS.md to dictate direction to generative AI tools. Both Claude and MuleSoft Vibes will automatically read the AGENTS.md at the start of a session.

# AGENTS.md — MuleSoft Development Project

## Project Overview
- **Project Name:** af-s-orderdb
- **API Name:** af-s-orderdb
- **Primary Use Cases:** System API connecting to the Order Database
- **Platform:** MuleSoft Anypoint Platform
- **Runtime:** Mule 4.9
- **Deployment Target:** CloudHub 2.0

## Notes for AI
- After loading this AGENTS.md file, also read and adhere to the
  XXXX-XXXX.md application specification file in the parent folder.
- If you cannot find this file, STOP and ASK.

Templates 

Templates are an excellent addition to AI generation for assets, such as applications and documentation. Examples of templates may be a sample of an existing application, or a Confluence page format for a design document. 

Code snippets and formats may exist as examples in the existing skills (remember the sub folders that obfuscate this low-level detail out of the main SKILL.md file). A template may be seen as a full version of something you wish the AI tool to replicate. 

To use a template, reference its location in the AGENTS.md file, along with the specification reference. Alternatively, we can use the template as a reference in the prompt, along with the skill to use, and any other asset you wish the AI tool to consider, such as the application specification. Say in the prompt or AGENTS.md file that the template takes precedence over the skill but report back on any anomalies.

Using the Skills: new project vs. existing project

Which skills you invoke, and in what order, differs depending on whether you’re starting from nothing or working against a codebase that already exists. The two flows share the same underlying skills; they just enter at different points.

New project – you have a specification, but no code: 

New project / specification ready
↓
mulesoft-code-generation
Build against the specification using the approved skills and prompts
↓
mulesoft-code-review
Check the result against current standards
↓
Gaps found?
If yes: iterate with human review
Remediate the code using prompts, then repeat the code-review step.
If no, or once resolved
↓
mulesoft-security-hardening
Run the final security-focused check
↓
Gaps found?
If yes: iterate with human review
Remediate the code using prompts, then repeat the security-hardening step.
If no, or once resolved
↓
mulesoft-munit-generation
Generate MUnit tests once the code is stable
↓
As-built documentation required?
If yes: mulesoft-documentor
Capture the final known-good state.
If no, or once documented
↓
Project baseline complete

Existing project -> you have code, but no guarantee it matches current standards. There’s nothing to generate, so the sequence starts at establishing ground truth instead:

Existing project / baseline unknown
↓
mulesoft-documentor
Understand what is currently there before making changes
↓
mulesoft-code-review
Identify code gaps against current standards
↓
Gaps identified?
If yes: iterate with human review
Remediate the code using prompts, then repeat the code-review step.
If no, or once resolved
↓
mulesoft-munit-review
Identify MUnit gaps against current standards
↓
Gaps identified?
If yes: iterate with human review
Remediate the MUnit tests using prompts, then repeat the MUnit-review step.
If no, or once resolved
↓
mulesoft-security-hardening
Review against the OWASP Top 10 and related security concerns
↓
Gaps identified?
If yes: iterate with human review
Remediate the code using prompts, then repeat the security-hardening step.
If no, or once resolved
↓
mulesoft-munit-generation
Fill test-coverage gaps once the code is in a corrected state
↓
As-built documentation needs refreshing?
If yes: mulesoft-documentor
Refresh the documentation to capture the corrected state.
If no, or once refreshed
↓
Baseline refreshed and standards aligned

Findings that measurably improved output 

  1. Ask the agent to replay its plan before generating code, and to surface clarifying questions first. Vibes in particular returns genuinely useful clarifying questions when prompted this way, rather than making silent assumptions. 
  2. Ground every new build in a written specification, not just a prompt. Specifications are good in markdown format as that is what the AI tool uses.  It will easily read .docx format but will use more tokens to convert and analyse. 
  3. Use a template if one exists, not just a prompt. If a template does not exist, consider creating one first, you will see how well the AI Tools can reproduce output that has the same look and feel, or consistency. 

What’s next

Generative AI tools are here to assist you in the project lifecycle but it’s important to remember, it can make mistakes. Internally, Adaptiv have adopted core lifecycle coverage (documentation, code generation, code review, MUnit generation, MUnit review and security hardening) in our development process, with a human-in-the-loop at all times, diligently reviewing and extensively testing generated output.

Our mission is to evolve our skills assets to reduce development time, while producing high quality, consistent, readable, secure and performant code for our clients. As we continue to invest in these capabilities ourselves, we encourage you to do the same. Understanding how to work effectively with AI tools is quickly becoming an important development skill, and one worth building now, you will be surprised at what they can do for you.

If you’re ready to go beyond the examples in this article, our MuleSoft AI Enablement offering can help you get started. Building on the same principles and skills demonstrated here, we provide the maintainable skills, guidance and guardrails needed to embed AI across your MuleSoft development lifecycle, without starting from scratch. Take the first step by booking a free session with our team to explore what AI-enabled MuleSoft development could look like for you.

Get in touch with our team, today