# Speech to Text With Speakers

## Links

- Product page URL: https://www.agentpmt.com/marketplace/speech-to-text-with-speakers
- Product markdown URL: https://www.agentpmt.com/marketplace/speech-to-text-with-speakers?format=agent-md
- Product JSON URL: https://www.agentpmt.com/marketplace/speech-to-text-with-speakers?format=agent-json

## Overview

- Product ID: 69ba14e4bbfb26a6333b14d3
- Type: function
- Unit type: request
- Price: 10000 credits
- Categories: AI & Machine Learning, Automation, Data Processing, Text Processing & Manipulation, Audio & Sound Design, Document Processing & OCR, Video & Streaming
- Generated at: 2026-08-19T04:10:51.307Z

### Page Description

Turn any audio recording into clean, searchable text in seconds. Transcribe voice memos, meetings, interviews, podcasts, and webinars with accurate speech recognition that handles accents and background noise. Get plain text for quick reference, SRT or WebVTT subtitles for video captioning, or rich JSON output with word-level timestamps and speaker identification. Choose from three tiers based on recording length — up to 15, 30, or 60 minutes — and optionally enable speaker diarization to label who said what, profanity filtering, and alternative transcripts for maximum accuracy.

### Agent Description

Transcribe audio from file_id or public_url with three tiered actions for recordings up to 15, 30, or 60 minutes. Transcribe actions return a task envelope with a task_id — short clips complete inline with status "completed", longer jobs are polled via get_task — with the finished transcript returned inline (text, SRT, VTT, or JSON) together with result_file_id and result_signed_url for the File Manager artifact, plus optional speaker diarization (up to 20 minutes), word timestamps, profanity filtering, and alternative transcriptions.

## Details

### Details

Turn any audio recording into clean, searchable text in seconds. Transcribe voice memos, meetings, interviews, podcasts, and webinars with accurate speech recognition that handles accents and background noise. Get plain text for quick reference, SRT or WebVTT subtitles for video captioning, or rich JSON output with word-level timestamps and speaker identification. Choose from three tiers based on recording length — up to 15, 30, or 60 minutes — and optionally enable speaker diarization to label who said what, profanity filtering, and alternative transcripts for maximum accuracy.

### Actions

- `transcribe_quick` (100 credits): Start an asynchronous transcription of audio up to 15 minutes. Returns a task_id immediately (short clips complete inline in the same response); poll get_task for the completed transcript and File Manager artifact.
- `transcribe_standard` (150 credits): Start an asynchronous transcription of audio up to 30 minutes. Returns a task_id immediately (short clips complete inline in the same response); poll get_task for the completed transcript and File Manager artifact.
- `transcribe_extended` (200 credits): Start an asynchronous transcription of audio up to 60 minutes. Returns a task_id immediately (short clips complete inline in the same response); poll get_task for the completed transcript and File Manager artifact.
- `get_task` (1 credits): Check a transcription task's progress and retrieve its result. Completed tasks carry the full transcription payload in outputs[0].
- `list_tasks` (1 credits): List recent transcription tasks with status and progress.

### Use Cases

Transcribe meeting recordings, Generate subtitles and captions for videos, Convert voice memos to searchable text, Transcribe podcast episodes, Create interview transcripts with speaker labels, Produce SRT or WebVTT subtitle files, Build searchable audio archives, Transcribe webinars and lectures, Analyze customer call recordings, Content repurposing from audio to text

### Workflows Using This Tool

#### Plaud Recordings to Google Drive Sync

Automatically backs up your Plaud recordings to Google Drive and keeps a tracking spreadsheet in Google Sheets. Each time you run this workflow, it downloads any new recordings from your Plaud account to a "Plaud Recordings" folder in Drive, creates a transcript for each recording (using Plaud's built-in transcripts when available, or automatic speech-to-text otherwise), saves the transcript alongside the audio file, identifies what each recording is about (meeting, interview, note, etc.), and logs everything in a spreadsheet with links to the audio and transcript files. You can run this as often as you like - it only processes new recordings and won't duplicate anything. Just connect your Plaud, Google Drive, and Google Sheets accounts and run it whenever you want to sync your latest recordings.

- Page URL: https://www.agentpmt.com/agent-workflow-skills/plaud-recordings-to-google-drive-sync
- Markdown URL: https://www.agentpmt.com/agent-workflow-skills/plaud-recordings-to-google-drive-sync?format=agent-md
- Published: 2026-08-07T14:43:24.041Z

#### Plaud Spoken Commitments to Telegram Reminders

Catches every date and time you commit to out loud in your Plaud recordings — "I'll send it Thursday", "meet Bob at 3" — and sends the whole batch to you on Telegram. Scans each new recording, pulls the transcript (reusing the recording's own transcript when one exists), extracts spoken commitments and resolves relative dates against the recording date and your timezone, and logs them to a "Plaud Commitments" Google Sheet. Every run then sends one Telegram digest listing all open commitments sorted by due date, with overdue items flagged. Mark a row "done" in the sheet to drop it from future digests. Schedule it daily and the digest becomes your morning commitments briefing — no phone number needed: connect once through the AgentPMT Telegram bot, and if you haven't yet, the run summary walks you through it. A processed-recordings tab guarantees each recording is scanned exactly once.

- Page URL: https://www.agentpmt.com/agent-workflow-skills/plaud-spoken-commitments-to-sms-reminders
- Markdown URL: https://www.agentpmt.com/agent-workflow-skills/plaud-spoken-commitments-to-sms-reminders?format=agent-md
- Published: 2026-07-22T14:01:25.974Z

#### Plaud Client Follow-Up Email Drafter

Drafts a follow-up email after every client call you record with Plaud — and never sends without your approval. Reviews each new recording, pulls the transcript (reusing the recording's own transcript when one exists), identifies genuine client conversations, and drafts a warm, concise follow-up recapping the discussion, what each side committed to, and the agreed next step — written strictly from what was said. Looks up the client's email address in your Google Contacts, then sends you the draft with the proposed recipient for approval; only after you approve (and confirm the address) is the email sent from your own Gmail account, so it lands in your Sent folder and replies come back to you. Every recording is logged in a Google Sheet so nothing is drafted twice. Skipped and rejected drafts are recorded with their reason.

- Page URL: https://www.agentpmt.com/agent-workflow-skills/plaud-client-follow-up-email-drafter
- Markdown URL: https://www.agentpmt.com/agent-workflow-skills/plaud-client-follow-up-email-drafter?format=agent-md
- Published: 2026-07-22T04:40:53.556Z

#### Plaud Meeting Recap to Slack

Posts a clean, speaker-labeled recap of every team meeting you record with Plaud into a Slack channel of your choice (defaults to #meeting-recaps). Reviews each new recording, pulls the transcript (reusing the recording's own transcript when one exists), identifies genuine multi-person team meetings — solo memos, client calls, and test recordings are skipped — and posts a formatted recap with participants, decisions, topics discussed, and action items with owners. A Google Sheet ledger guarantees each recording is handled exactly once, so the workflow can run on a schedule without duplicate posts. The Slack app must be a member of the target channel.

- Page URL: https://www.agentpmt.com/agent-workflow-skills/plaud-meeting-recap-to-slack
- Markdown URL: https://www.agentpmt.com/agent-workflow-skills/plaud-meeting-recap-to-slack?format=agent-md
- Published: 2026-07-22T04:25:40.501Z

#### Plaud Master Action-Item Tracker

Collects every action item spoken in your Plaud recordings into one master Google Sheet. Scans each new recording, pulls its transcript (reusing the recording's own transcript when one exists) and AI summary, extracts concrete tasks with the responsible owner and any spoken due date (resolved to a real calendar date from the recording date), and appends them to an "Action Items" tab with status "open" — deduplicated within and across meetings so the same commitment never appears twice. A "Processed Recordings" tab guarantees each recording is scanned exactly once, so the workflow can run on a schedule. Solves the problem of action items being trapped inside individual meeting summaries.

- Page URL: https://www.agentpmt.com/agent-workflow-skills/plaud-master-action-item-tracker
- Markdown URL: https://www.agentpmt.com/agent-workflow-skills/plaud-master-action-item-tracker?format=agent-md
- Published: 2026-07-22T04:22:11.630Z

#### Plaud Call Logger for Pipedrive

Turns your Plaud voice recordings of client calls into Pipedrive records automatically. Reviews every new recording, identifies genuine client conversations, transcribes them (reusing the recording's own transcript when one exists), saves the transcript to a "Plaud Call Logs" Drive folder, matches the spoken client name and company to a Pipedrive person and open deal, attaches a call note with the discussion summary, commitments, next steps, and transcript link, and creates a follow-up activity dated from spoken commitments. Keeps a Google Sheet ledger so every recording is processed exactly once; recordings that are not client calls are logged and skipped. Re-runnable and idempotent.

- Page URL: https://www.agentpmt.com/agent-workflow-skills/plaud-call-logger-for-pipedrive
- Markdown URL: https://www.agentpmt.com/agent-workflow-skills/plaud-call-logger-for-pipedrive?format=agent-md
- Published: 2026-07-22T04:22:05.790Z

### Related Content

#### Artificial Intelligence Medical Scribe, Captions, and Transcripts on AgentPMT

- Type: article
- Page URL: https://www.agentpmt.com/articles/artificial-intelligence-medical-scribe-captions-and-transcripts-on-agentpmt
- Markdown URL: https://www.agentpmt.com/articles/artificial-intelligence-medical-scribe-captions-and-transcripts-on-agentpmt?format=agent-md
Speech to Text With Speakers, built by Apoth3osis, is now live on AgentPMT, a managed connector that turns any recording into accurate text, SRT/VTT captions, or timestamped JSON with speaker diarization across 15-, 30-, and 60-minute tiers. Agents call it through the dynamic MCP server and pay only when a transcription succeeds.

#### Animal Artificial Intelligence Learns to Read the Wild

- Type: article
- Page URL: https://www.agentpmt.com/articles/animal-artificial-intelligence-learns-to-read-the-wild
- Markdown URL: https://www.agentpmt.com/articles/animal-artificial-intelligence-learns-to-read-the-wild?format=agent-md
In a single week, a cluster of research releases showed AI in the animal world moving from finding and counting animals to reading them: re-identifying individuals on GPU-free hardware, inferring diet from feeding sounds, and mapping a songbird's calls. With the capture problem largely solved, the advantage now shifts to the operational work around the model, choosing it on cost and quality, keeping a human on high-stakes calls, and recording why it decided what it did.

## Documentation

No platform documentation is currently linked to this product.

## Integration Details

### DynamicMCP

- Setup page URL: https://www.agentpmt.com/dynamic-mcp
- Claude setup guide: https://www.agentpmt.com/dynamic-mcp#platform=claude
- ChatGPT setup guide: https://www.agentpmt.com/dynamic-mcp#platform=chatgpt
- Cursor setup guide: https://www.agentpmt.com/dynamic-mcp#platform=cursor
- Windsurf setup guide: https://www.agentpmt.com/dynamic-mcp#platform=windsurf

Use the local router for command-based MCP clients. It forwards requests to `https://api.agentpmt.com/mcp` and does not execute tools locally.

```bash
npm install -g @agentpmt/mcp-router
agentpmt-setup
```

### REST API

The live page renders cURL, Python, JavaScript, and Node.js examples. Logged-in users see those examples prefilled with their own API and budget credentials.

- Purchase endpoint: https://api.agentpmt.com/products/purchase
- Authorization format: `Bearer <base64(apiKey:budgetKey)>`

```bash
curl -X POST "https://api.agentpmt.com/products/purchase" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer eW91ci1hcGkta2V5LWhlcmU6eW91ci1idWRnZXQta2V5LWhlcmU=" \
  -d '{
    "product_id": "69ba14e4bbfb26a6333b14d3",
    "parameters": {
      "action": "transcribe_quick",
      "output_format": "text",
      "enable_diarization": false,
      "enable_word_timestamps": false,
      "remove_filler_words": true,
      "enable_profanity_filter": false,
      "max_alternatives": 1
    }
  }'
```

### Autonomous Agents

Autonomous agents can access this tool through AgentAddress credit balances or direct x402 payments. Use the Autonomous Agent API reference for endpoint shapes after choosing the access pattern below.

- Autonomous Agent API reference URL: https://www.agentpmt.com/docs/api-reference/autonomous-agents
- Autonomous Agent API reference markdown URL: https://www.agentpmt.com/docs/api-reference/autonomous-agents?format=agent-md
- Credit-Based Access Using AgentAddress: https://www.agentpmt.com/docs/autonomous-agents/credit-based-tool-usage-with-agentaddress
- AgentAddress is preferred for persistent file access, stored platform state, and maximum tool use ability across repeated calls.
- Direct x402 is for independent one-off tool calls that do not require shared files or stored platform state.
- Direct x402 public payments: USDC on Base, Arbitrum, Optimism, Polygon, and Avalanche.

#### Product Skill Package

This product has a published Agent Skill package for product-specific operating instructions.

- Skill slug: speech-to-text-with-speakers
- Version: 1.0.6
- Download SKILL.md: https://raw.githubusercontent.com/AgentPMT/agent-skills/main/skills/speech-to-text-with-speakers/SKILL.md
- Package source: https://github.com/AgentPMT/agent-skills/tree/main/skills/speech-to-text-with-speakers
- OpenClaw listing: https://clawhub.ai/agentpmt/speech-to-text-with-speakers
- OpenClaw install: `openclaw skills install speech-to-text-with-speakers`
- skills.sh install: `npx skills add AgentPMT/agent-skills --skill speech-to-text-with-speakers`
- Last published: 2026-07-22T00:25:53.259Z

### Schema

#### Parameters

- Schema type: actions

```json
{
  "actions": {
    "transcribe_quick": {
      "description": "Start an asynchronous transcription of audio up to 15 minutes. Returns a task_id immediately (short clips complete inline in the same response); poll get_task for the completed transcript and File Manager artifact.",
      "properties": {
        "file_id": {
          "type": "string",
          "description": "File ID from a prior upload. Provide either file_id or public_url.",
          "required": false
        },
        "public_url": {
          "type": "string",
          "description": "HTTPS URL to a downloadable audio file. Provide either public_url or file_id.",
          "required": false
        },
        "language_code": {
          "type": "string",
          "description": "Optional BCP-47 language code such as en-US; defaults to en-US if omitted.",
          "required": false
        },
        "output_format": {
          "type": "string",
          "description": "Output format for the transcription result. The completed task's outputs include the selected content inline plus result_file_id and result_signed_url for the File Manager artifact.",
          "required": false,
          "enum": [
            "text",
            "srt",
            "vtt",
            "json"
          ],
          "default": "text"
        },
        "enable_diarization": {
          "type": "boolean",
          "description": "Enable speaker diarization when supported by the audio and model. Not supported when remove_filler_words is false.",
          "required": false,
          "default": false
        },
        "enable_word_timestamps": {
          "type": "boolean",
          "description": "Include word-level timing data in the output.",
          "required": false,
          "default": false
        },
        "remove_filler_words": {
          "type": "boolean",
          "description": "When true (default), return a cleaned transcript with disfluencies removed. When false, preserve filler words and disfluencies; this path does not support diarization or max_alternatives greater than 1.",
          "required": false,
          "default": true
        },
        "enable_profanity_filter": {
          "type": "boolean",
          "description": "Mask profanity in the returned transcript.",
          "required": false,
          "default": false
        },
        "max_alternatives": {
          "type": "integer",
          "description": "Maximum number of alternative transcripts to return. Must be 1 when remove_filler_words is false.",
          "required": false,
          "minimum": 1,
          "maximum": 5,
          "default": 1
        }
      },
      "price_per_unit": 100
    },
    "transcribe_standard": {
      "description": "Start an asynchronous transcription of audio up to 30 minutes. Returns a task_id immediately (short clips complete inline in the same response); poll get_task for the completed transcript and File Manager artifact.",
      "properties": {
        "file_id": {
          "type": "string",
          "description": "File ID from a prior upload. Provide either file_id or public_url.",
          "required": false
        },
        "public_url": {
          "type": "string",
          "description": "HTTPS URL to a downloadable audio file. Provide either public_url or file_id.",
          "required": false
        },
        "language_code": {
          "type": "string",
          "description": "Optional BCP-47 language code such as en-US; defaults to en-US if omitted.",
          "required": false
        },
        "output_format": {
          "type": "string",
          "description": "Output format for the transcription result. The completed task's outputs include the selected content inline plus result_file_id and result_signed_url for the File Manager artifact.",
          "required": false,
          "enum": [
            "text",
            "srt",
            "vtt",
            "json"
          ],
          "default": "text"
        },
        "enable_diarization": {
          "type": "boolean",
          "description": "Enable speaker diarization when supported by the audio and model. Not supported when remove_filler_words is false.",
          "required": false,
          "default": false
        },
        "enable_word_timestamps": {
          "type": "boolean",
          "description": "Include word-level timing data in the output.",
          "required": false,
          "default": false
        },
        "remove_filler_words": {
          "type": "boolean",
          "description": "When true (default), return a cleaned transcript with disfluencies removed. When false, preserve filler words and disfluencies; this path does not support diarization or max_alternatives greater than 1.",
          "required": false,
          "default": true
        },
        "enable_profanity_filter": {
          "type": "boolean",
          "description": "Mask profanity in the returned transcript.",
          "required": false,
          "default": false
        },
        "max_alternatives": {
          "type": "integer",
          "description": "Maximum number of alternative transcripts to return. Must be 1 when remove_filler_words is false.",
          "required": false,
          "minimum": 1,
          "maximum": 5,
          "default": 1
        }
      },
      "price_per_unit": 150
    },
    "transcribe_extended": {
      "description": "Start an asynchronous transcription of audio up to 60 minutes. Returns a task_id immediately (short clips complete inline in the same response); poll get_task for the completed transcript and File Manager artifact.",
      "properties": {
        "file_id": {
          "type": "string",
          "description": "File ID from a prior upload. Provide either file_id or public_url.",
          "required": false
        },
        "public_url": {
          "type": "string",
          "description": "HTTPS URL to a downloadable audio file. Provide either public_url or file_id.",
          "required": false
        },
        "language_code": {
          "type": "string",
          "description": "Optional BCP-47 language code such as en-US; defaults to en-US if omitted.",
          "required": false
        },
        "output_format": {
          "type": "string",
          "description": "Output format for the transcription result. The completed task's outputs include the selected content inline plus result_file_id and result_signed_url for the File Manager artifact.",
          "required": false,
          "enum": [
            "text",
            "srt",
            "vtt",
            "json"
          ],
          "default": "text"
        },
        "enable_diarization": {
          "type": "boolean",
          "description": "Enable speaker diarization when supported by the audio and model. Not supported when remove_filler_words is false.",
          "required": false,
          "default": false
        },
        "enable_word_timestamps": {
          "type": "boolean",
          "description": "Include word-level timing data in the output.",
          "required": false,
          "default": false
        },
        "remove_filler_words": {
          "type": "boolean",
          "description": "When true (default), return a cleaned transcript with disfluencies removed. When false, preserve filler words and disfluencies; this path does not support diarization or max_alternatives greater than 1.",
          "required": false,
          "default": true
        },
        "enable_profanity_filter": {
          "type": "boolean",
          "description": "Mask profanity in the returned transcript.",
          "required": false,
          "default": false
        },
        "max_alternatives": {
          "type": "integer",
          "description": "Maximum number of alternative transcripts to return. Must be 1 when remove_filler_words is false.",
          "required": false,
          "minimum": 1,
          "maximum": 5,
          "default": 1
        }
      },
      "price_per_unit": 200
    },
    "get_task": {
      "description": "Check a transcription task's progress and retrieve its result. Completed tasks carry the full transcription payload in outputs[0].",
      "properties": {
        "task_id": {
          "type": "string",
          "description": "Task ID returned by a transcribe action.",
          "required": true
        }
      },
      "price_per_unit": 1
    },
    "list_tasks": {
      "description": "List recent transcription tasks with status and progress.",
      "properties": {
        "limit": {
          "type": "integer",
          "description": "Maximum number of tasks to return.",
          "required": false,
          "default": 20,
          "minimum": 1,
          "maximum": 100
        }
      },
      "price_per_unit": 1
    }
  }
}
```

### Usage Instructions

# Speech to Text

Transcribe audio with one tool. Transcribe actions run as background tasks: the submit response returns a `task_id` immediately, and short clips usually complete inline in that same response. Poll `get_task` for anything still processing.

## Tool Call Format

```json
{
  "action": "get_instructions"
}
```

```json
{
  "action": "transcribe_quick",
  "file_id": "FILE_ID",
  "language_code": "en-US",
  "output_format": "text"
}
```

```json
{
  "action": "transcribe_standard",
  "public_url": "https://example.com/meeting.m4a",
  "output_format": "vtt",
  "enable_word_timestamps": true,
  "enable_diarization": true
}
```

```json
{
  "action": "transcribe_extended",
  "public_url": "https://example.com/interview.webm",
  "output_format": "json",
  "max_alternatives": 2
}
```

```json
{
  "action": "get_task",
  "task_id": "TASK_ID"
}
```

```json
{
  "action": "list_tasks",
  "limit": 20
}
```

## Actions

- `transcribe_quick`: audio up to 15 minutes. Price: 100 credits.
- `transcribe_standard`: audio up to 30 minutes. Price: 150 credits.
- `transcribe_extended`: audio up to 60 minutes. Price: 200 credits.
- `get_task`: check a transcription task's progress and retrieve its result. Price: 1 credit.
- `list_tasks`: list recent transcription tasks. Price: 1 credit.

## Async task flow

- Every transcribe action returns a task envelope: `{action, task_id, status, ...}`. ALWAYS check `status` before polling — short clips finish within the submit request and return `status: "completed"` with `outputs` inline, costing zero polls.
- When `status` is `"processing"`, poll `get_task` with the returned `task_id` every 10-15 seconds. `progress` advances as the job moves through download, validation, and recognition.
- A completed task carries the full transcription payload in `outputs[0]`: the selected content inline (`text`, `srt_content`, `vtt_content`, or `json_data`), plus `speakers`, `word_count`, `confidence_score`, `audio_metadata`, and the File Manager artifact fields `result_file_id`/`result_signed_url`.
- A failed task carries `error`. Tier-limit failures also carry `error_details` with `recommended_actions` — resubmit with the suggested larger tier.
- If the service restarts mid-transcription, in-flight Google batch jobs resume automatically. Jobs that cannot resume are marked failed with an instruction to resubmit; they never sit in `processing` forever.

## Billing

- Transcribe actions charge on submission. A task that later fails (for example, audio too long for the tier) is not auto-refunded — pick the tier from the recording length you already know, and check `error_details.recommended_actions` before resubmitting.

## Notes

- Provide either `file_id` or `public_url`.
- `public_url` must be an HTTPS URL and cannot point to private or internal network addresses.
- Audio downloads are capped at 150MB. If a recording is larger, compress it to mp3 or m4a and retry.
- If `language_code` is omitted, the tool defaults to `en-US`.
- Supported output formats: `text`, `srt`, `vtt`, `json`.
- Optional controls: `enable_diarization`, `enable_word_timestamps`, `remove_filler_words`, `enable_profanity_filter`, `max_alternatives`.
- `remove_filler_words` defaults to `true`, which uses Google STT V2's cleaned transcript path.
- Set `remove_filler_words` to `false` to preserve disfluencies through Vercel AI Gateway using the `openai/whisper-1` gateway model slug. This path always requests word-level timestamps from the gateway for clipping workflows.
- `remove_filler_words=false` does not support `enable_diarization=true` or `max_alternatives` greater than `1`; use the default cleaned path for those features.
- `enable_diarization=true` supports audio up to 20 minutes (a provider limit). Longer recordings fail with guidance: disable diarization, or split the audio into 20-minute segments and transcribe them individually.
- With `enable_diarization=true`, `text` output is speaker-labelled one line per turn (`[0:04] Speaker 0: ...`), inline and in the stored `transcription.txt`, so it is readable without reformatting. Without diarization it is the flat transcript. Speaker numbers match the `speaker_tag` values in `speakers`/`json_data`.
- During invocations with File Manager storage available, `text` and `json` results are stored for every successful transcription; `srt` and `vtt` results are stored when the generated subtitle content is non-empty. The corresponding filenames are `transcription.txt`, `transcription.json`, `transcription.srt`, and `transcription.vtt`.
- Stored artifacts are returned through the `result_file_id` and `result_signed_url` fields in `outputs[0]`. Agents should use `result_file_id` for later File Manager operations.
- If transcription succeeds but artifact storage fails, the inline result remains available and `result_file_error` explains that the File Manager file could not be created.

### Frequently Asked Questions

#### How do I connect this tool to an external agent?

- Page URL: https://www.agentpmt.com/faq
- Markdown URL: https://www.agentpmt.com/faq?format=agent-md

You can install the local MCP server by opening a terminal and running:

```
npm install -g @agentpmt/mcp-router
agentpmt-setup
```

This will connect you to local agents like Claude Code, Windsurf, Grok Build, Cursor, etc.

Alternatively you can connect to the hosted version with this config block, no installation required:

```
{
  "mcpServers": {
    "agentpmt": {
      "type": "streamable-http",
      "url": "https://api.agentpmt.com/mcp",
      "headers": {
        "Authorization": "Bearer <AGENTPMT_BEARER_TOKEN>",
        "x-instance-metadata": "{\"client\":\"generic-mcp\",\"platform\":\"remote\"}"
      }
    }
  }
}
```

[View MCP Connection Instructions](/docs/mcp-reference/connection) for more details.

#### How does an external agent use this tool?

- Page URL: https://www.agentpmt.com/faq
- Markdown URL: https://www.agentpmt.com/faq?format=agent-md

After the external agent is connected to an Agent Group that can use this tool, paste this prompt into the agent:

> Use the AgentPMT-Tool-Search-and-Execution tool. First call action 'get\_instructions' so you know how to use the tool search interface. Then call action 'get\_schema' with tool\_id 69ba14e4bbfb26a6333b14d3 ("Speech to Text With Speakers"). After reading the schema and any returned instructions, tell me what this tool can do, we are going to be using it

The agent should fetch the tool schema first, collect the required parameters for your request, and then call the tool through AgentPMT.

### Dependencies

This product has no public dependency products.