Serveur MCP

BananaBanana Image, Video & Speech Generation

pro.bananabanana/image-video
Design et création Médias et contenu Public et accessible MCP 2026-07-28

Ce que fait ce MCP

Generates and edits images, videos, and speech using hosted AI media models, with job tracking, model selection, pricing, and account balance tools.

edit_image
Edit / refine a previously generated image with a text instruction on Nano Banana 2 Lite / 2 / Pro (multi-turn editing: change colors, remove objects, restyle, etc.). This tool modifies an image that already exists — use generate_image to create a new image from a prompt. Pass the job_id of a COMPLETED image generation as source_generation_id. Charged like a single image of the chosen model/resolution; auto-refund on failure. Example: {"source_generation_id": "cmxyz...", "prompt": "make the background pure white and add soft shadow"}
Schéma d’entrée
{'type': 'object', 'required': ['source_generation_id', 'prompt'], 'properties': {'seed': {'type': 'integer', 'maximum': 2147483647, 'minimum': 0}, 'model': {'enum': ['nano-banana-2-lite', 'nano-banana-2', 'nano-banana-pro'], 'type': 'string', 'default': 'nano-banana-2'}, 'prompt': {'type': 'string', 'maxLength': 32000, 'description': 'The edit instruction (up to 32000 characters).'}, 'resolution': {'enum': ['512', '1024', '2048', '4096'], 'type': 'string', 'default': '1024'}, 'aspect_ratio': {'enum': ['1:1', '3:2', '2:3', '4:3', '3:4', '4:5', '5:4', '9:16', '16:9', '21:9'], 'type': 'string', 'default': '1:1'}, 'output_format': {'enum': ['jpeg', 'png', 'webp'], 'type': 'string', 'default': 'jpeg'}, 'relaxed_filter': {'type': 'boolean', 'default': True, 'description': 'Enabled by default for Google images. Uses safetySettings thresholds OFF (no configurable threshold blocking) and personGeneration=ALLOW_ADULT. Pass false to use BLOCK_ONLY_HIGH and omit personGeneration. Runs only on our own Google keys: when none is free, an explicit true is refused with SERVICE_UNAVAILABLE before any charge, and the default silently falls back to the standard filter (get_result then reports relaxed_filter: false). Ignored by other image providers and by every video model. OFF and BLOCK_NONE are image/Gemini safety options, not video parameters. Non-configurable checks still apply, including checks on generated images and input media. Celebrity refusals (support codes 15236754 / 29310472) are not disabled by this flag. get_result explains the category and stage when Google supplies support codes; failed jobs are refunded.'}, 'idempotency_key': {'type': 'string', 'maxLength': 64}, 'source_generation_id': {'type': 'string', 'description': 'job_id of a completed image generation owned by this account.'}}, 'additionalProperties': False}
edit_video
Edit an EXISTING video (video-to-video). Default model omni-flash (Gemini Omni Flash): restyle it, replace or add objects, relight the scene, change the mood — motion and composition of the source clip are preserved. Billed by output length at $0.10/s; cost confirmation is mandatory (first call returns the quote and charges nothing). With model "wan-3.0" (Alibaba Wan 3.0) the source is read as a reference (up to 15 s): mode "edit" rebuilds the scene as the prompt says (framing follows the source, you pick the output length and resolution), mode "extend" shoots the next scene with the same characters and setting — a story continuation, not a frame-exact extension. Wan bills the SOURCE SECONDS TOGETHER WITH THE OUTPUT (the provider charges for both) and source + output must fit in 30 s. Source: either source_generation_id (a completed video from this account — see list_generations) or video_url (public http(s) link, max 200 MB). The source is normalised to MP4 720p and, for mode "edit", the FIRST 10 SECONDS (model limit); output is 720p with sound, aspect ratio follows the source. OUTPUT LENGTH ALWAYS EQUALS SOURCE LENGTH (the model cannot stretch or shorten a clip), so duration only works downwards: it trims the source to the first N seconds. Omit duration to edit the whole clip — the quote tells you the resolved length. mode: "extend" instead appends a continuation to the end of the clip — duration is then the LENGTH OF THE APPENDED PART (3–10 s), the source plays first and the total is capped at 40 s. Extension reads the WHOLE source (up to 37 s), so extensions chain: extend the result again until 40 s. Extension is billed on the full resulting length. With mode "extend", reference_images can bring a subject or product into the continuation ("the character from the reference image enters the scene"). Returns a job_id; poll get_result. Failed edits are auto-refunded. Example: {"prompt": "make the whole scene look like a pencil sketch, keep the motion identical", "source_generation_id": "clx…", "confirm_cost": 1}
Schéma d’entrée
{'type': 'object', 'required': ['prompt'], 'properties': {'mode': {'enum': ['edit', 'extend'], 'type': 'string', 'default': 'edit', 'description': '"edit" rewrites the clip; "extend" continues it. omni-flash: edit keeps the length, extend appends a 3–10 s continuation (total capped at 40 s). wan-3.0: edit rebuilds the scene at the chosen duration, extend shoots the next scene.'}, 'model': {'enum': ['omni-flash', 'omni-flash-1.0', 'wan-3.0'], 'type': 'string', 'default': 'omni-flash', 'description': 'omni-flash (Gemini Omni 1.1) keeps the motion and length of the source and can extend the scene; omni-flash-1.0 is the previous generation (edit only, same length); wan-3.0 reads the source as a reference and renders a new clip of the chosen duration (edit or extend by prompt).'}, 'prompt': {'type': 'string', 'maxLength': 32000, 'description': 'What to change in the video (up to 20000 characters).'}, 'duration': {'enum': [3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30], 'type': 'integer', 'description': 'omni-flash with mode "edit": trim the source to the first N seconds and edit only that part; values above the source length are ignored. With mode "extend": the length of the appended continuation (3–10 s). wan-3.0: the length of the OUTPUT clip — 4, 6, 8, 10, 15, 20, 30 s (default 8); source + output must fit in 30 s.'}, 'video_url': {'type': 'string', 'description': 'Public http(s) URL of the source video (mp4/mov/webm/mkv/avi/wmv/flv/3gpp, max 200 MB). Use instead of source_generation_id.'}, 'resolution': {'enum': ['480p', '720p', '1080p'], 'type': 'string', 'description': 'wan-3.0 only: output resolution (default 720p). omni-flash edits are always 720p.'}, 'source_ref': {'type': 'string', 'description': 'Returned by the quote when video_url is used: pass it back with confirm_cost to reuse the already downloaded clip instead of downloading it again.'}, 'with_audio': {'type': 'boolean', 'default': True, 'description': 'wan-3.0 only: sound is on by default and free — pass false for a silent clip. omni-flash always generates audio.'}, 'audio_prompt': {'type': 'string', 'maxLength': 500, 'description': 'Describe the desired sound — Omni Flash always generates audio.'}, 'confirm_cost': {'type': 'number', 'description': 'The quoted USD cost you accept. Omit on the first call to get the quote.'}, 'idempotency_key': {'type': 'string', 'maxLength': 64, 'description': 'Optional unique key; retries with the same key never double-charge.'}, 'reference_images': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 10, 'description': 'omni-flash with mode "extend" only: images of subjects, products or characters the continuation should bring into the scene — a completed image job_id, a public http(s) URL, or a data:image/...;base64,... value (≤10 MB). Refer to them in the prompt ("the item from the reference image enters the frame"). Ignored by mode "edit" and by wan-3.0.'}, 'source_generation_id': {'type': 'string', 'description': 'job_id of a completed video generation on this account.'}}, 'additionalProperties': False}
generate_image
Start an AI image generation (Google Nano Banana family, OpenAI GPT Image 2.5, Alibaba Qwen-Image 3.0 Pro). Charges the account balance immediately and returns a job_id — poll get_result for the finished image URLs. Typical completion: 10–60 seconds. The default nano-banana-2-lite model is the cheapest Google option and produces 1024px images; choose nano-banana-2 explicitly for 512px, 2048px or 4096px output. Optional reference_images provide the model with the actual subject, product, character or style pixels; Nano Banana Pro supports up to 14 references (per-model limits in list_models). Costs $0.03–$0.20 per image depending on model and resolution (see list_models). Failed generations are automatically refunded. relaxed_filter is true by default for Google images; pass false to opt out. Non-configurable checks still apply. Generating several images at once (number_of_images > 1) is a batch: the first call returns a price quote and charges nothing — repeat the call with confirm_cost set to the quoted amount to start. Example: {"prompt": "studio photo of a ceramic mug on linen, soft daylight", "model": "nano-banana-2-lite", "aspect_ratio": "4:5", "resolution": "1024"}
Schéma d’entrée
{'type': 'object', 'required': ['prompt'], 'properties': {'seed': {'type': 'integer', 'maximum': 2147483647, 'minimum': 0, 'description': 'For reproducible results.'}, 'model': {'enum': ['nano-banana-2-lite', 'nano-banana-2', 'nano-banana-pro', 'gpt-image-2.5-flare', 'gpt-image-2.5-sunburst', 'qwen-image-3.0-pro'], 'type': 'string', 'default': 'nano-banana-2-lite', 'description': 'nano-banana-2-lite: cheapest Google default, 1024 only. nano-banana-2: choose for 512, 2048 or 4096 output. nano-banana-pro: top Google quality, up to 4K. gpt-image-2.5-flare / gpt-image-2.5-sunburst: OpenAI — strongest at readable in-image text and long literal briefs; 1024, 2048 or 4096 (4K = 3840 on the long side); Flare is the fast 2.5 tier, Sunburst the precision tier; they ignore seed and relaxed_filter, and with reference_images they take the orientation of the first reference (omit aspect_ratio there, it is rejected). qwen-image-3.0-pro: Alibaba — crisp small text, dense layouts (posters, menus, UI); 1024 or 2048 only, up to 3 references, seed + negative_prompt. Cheapest per image: nano-banana-2-lite at $0.03.'}, 'prompt': {'type': 'string', 'maxLength': 32000, 'description': 'What to generate. English works best. Length limit depends on the model — 32000 characters on every image model here; list_models reports max_prompt_chars for each.'}, 'resolution': {'enum': ['512', '1024', '2048', '4096'], 'type': 'string', 'default': '1024', 'description': 'Lite requires 1024. Choose nano-banana-2 for 512, 2048 or 4096; 4096 costs the most. GPT Image models have no 512; qwen-image-3.0-pro accepts 1024 or 2048 only.'}, 'aspect_ratio': {'enum': ['1:1', '3:2', '2:3', '4:3', '3:4', '4:5', '5:4', '9:16', '16:9', '21:9'], 'type': 'string', 'default': '1:1'}, 'confirm_cost': {'type': 'number', 'description': 'Required for batches (number_of_images > 1): the quoted total USD cost you accept.'}, 'output_format': {'enum': ['jpeg', 'png', 'webp'], 'type': 'string', 'default': 'jpeg'}, 'relaxed_filter': {'type': 'boolean', 'default': True, 'description': 'Enabled by default for Google images. Uses safetySettings thresholds OFF (no configurable threshold blocking) and personGeneration=ALLOW_ADULT. Pass false to use BLOCK_ONLY_HIGH and omit personGeneration. Runs only on our own Google keys: when none is free, an explicit true is refused with SERVICE_UNAVAILABLE before any charge, and the default silently falls back to the standard filter (get_result then reports relaxed_filter: false). Ignored by other image providers and by every video model. OFF and BLOCK_NONE are image/Gemini safety options, not video parameters. Non-configurable checks still apply, including checks on generated images and input media. Celebrity refusals (support codes 15236754 / 29310472) are not disabled by this flag. get_result explains the category and stage when Google supplies support codes; failed jobs are refunded.'}, 'idempotency_key': {'type': 'string', 'maxLength': 64, 'description': 'Optional unique key; retries with the same key never double-charge.'}, 'negative_prompt': {'type': 'string', 'maxLength': 1000}, 'number_of_images': {'type': 'integer', 'default': 1, 'maximum': 4, 'minimum': 1, 'description': 'Variants per call. >1 requires confirm_cost.'}, 'reference_images': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 14, 'description': "Actual visual references for the generated image. Each item is a job_id of a completed image generation on this account, a public http(s) image URL, or an inline data:image/png|jpeg|webp;base64,... URL (max 10 MB each). For a local file, read and base64-encode its bytes into a data URL; a bare local path such as C:\\\\photo.jpg cannot be read by this remote server. Google and Qwen models keep aspect_ratio with references; GPT Image models follow the first reference's orientation instead and reject an explicit aspect_ratio."}}, 'additionalProperties': False}
generate_speech
Generate natural speech with Gemini 3.1 Flash TTS Preview. This synchronous tool returns a hosted WAV URL directly (no get_result polling). Supports one voice or an exactly two-speaker dialogue, automatic language detection or a BCP-47 language_code, natural-language direction for accent/tone/pace, and inline performance tags such as [whispers], [laughs], [very slow] and [excited]. Price is $0.01 per started 200 transcript characters; the account is charged only after Google has returned valid audio. Example: {"text":"[cheerfully] Welcome to BananaBanana!","voice":"Kore","style":"Warm product announcement, medium pace."}
Schéma d’entrée
{'type': 'object', 'required': ['text'], 'properties': {'text': {'type': 'string', 'maxLength': 6000, 'description': 'Exact transcript to speak. For dialogue, prefix every turn with the matching speaker name, e.g. Sam: Hello.'}, 'style': {'type': 'string', 'maxLength': 1500, 'description': 'Optional overall direction: persona, scene, emotion, accent, pace, pronunciation and delivery notes.'}, 'voice': {'enum': ['Achernar', 'Achird', 'Algenib', 'Algieba', 'Alnilam', 'Aoede', 'Autonoe', 'Callirrhoe', 'Charon', 'Despina', 'Enceladus', 'Erinome', 'Fenrir', 'Gacrux', 'Iapetus', 'Kore', 'Laomedeia', 'Leda', 'Orus', 'Pulcherrima', 'Puck', 'Rasalgethi', 'Sadachbia', 'Sadaltager', 'Schedar', 'Sulafat', 'Umbriel', 'Vindemiatrix', 'Zephyr', 'Zubenelgenubi'], 'type': 'string', 'default': 'Kore', 'description': 'Single-speaker voice. Ignored when speakers is provided.'}, 'speakers': {'type': 'array', 'items': {'type': 'object', 'required': ['name', 'voice'], 'properties': {'name': {'type': 'string', 'maxLength': 40, 'minLength': 1}, 'voice': {'enum': ['Achernar', 'Achird', 'Algenib', 'Algieba', 'Alnilam', 'Aoede', 'Autonoe', 'Callirrhoe', 'Charon', 'Despina', 'Enceladus', 'Erinome', 'Fenrir', 'Gacrux', 'Iapetus', 'Kore', 'Laomedeia', 'Leda', 'Orus', 'Pulcherrima', 'Puck', 'Rasalgethi', 'Sadachbia', 'Sadaltager', 'Schedar', 'Sulafat', 'Umbriel', 'Vindemiatrix', 'Zephyr', 'Zubenelgenubi'], 'type': 'string'}}, 'additionalProperties': False}, 'maxItems': 2, 'minItems': 2, 'description': 'Exactly two dialogue speakers. Their names must prefix the turns in text.'}, 'language_code': {'type': 'string', 'description': 'Optional language/locale such as en-US, ru-RU, ja-JP or es-MX. Omit for automatic detection.'}, 'idempotency_key': {'type': 'string', 'maxLength': 64, 'description': 'Optional unique key; retries with the same key never double-charge.'}}, 'additionalProperties': False}
generate_video
Start an AI video generation (Google Veo 3.1 family, Gemini Omni Flash, Alibaba Wan 3.0 or xAI Grok Imagine Video 1.5). EXPENSIVE: $0.10–$6.00 per clip. Cost confirmation is mandatory: the first call always returns a USD quote and charges nothing — repeat the call with confirm_cost set to the quoted amount to actually start. Returns a job_id; poll get_result (videos take 1–10+ minutes). Failed generations are auto-refunded. Models: veo-3.1-fast (default, good quality/price), veo-3.1 (best Veo quality), veo-3.1-lite (cheapest, 720p/1080p), wan-3.0 (Alibaba Wan 3.0 — clips up to 30 s, 480p/720p/1080p, optional audio, first and last frame, seed, and VIDEO INPUTS via reference_videos: edit a clip, continue its story or carry its characters into a new scene, all driven by the prompt), grok-imagine-video-1.5 (xAI Grok Imagine Video 1.5 — 4–15 s at any whole second, 720p/1080p, native sound always on, priced per second by resolution — $0.14/s at 720p, $0.25/s at 1080p, input images free — aspect ratios 16:9 / 9:16 / 1:1 / 3:2 / 2:3, a first_frame at both 720p and 1080p, and up to 7 reference_images at 720p only — 1080p takes a first_frame but no references; no seed, last frame or video input), omni-flash (Gemini Omni 1.1 Flash — always has sound, 3–10 s, priced per second by resolution — $0.03/s at 360p up to $0.30/s at 4k; supports a last frame for interpolation, conversational editing via edit_from_generation_id, and scene extension via edit_video mode=extend), omni-flash-1.0 (the previous Omni generation, kept as an option: 720p at $0.10/s or 1080p upscaled at $0.12/s, 4/6/8/10 s, conversational editing, no last frame / scene extension / video references). Content filtering differs by model: omni-flash is by far the strictest, so a clip rejected as SAFETY_FILTERED on omni-flash is often produced by veo-3.1-fast without changing a word. Video has no configurable safety settings on Google's side (relaxed_filter is accepted for Veo but only pins its default personGeneration=allow_adult), and refused clips are refunded, so a retry is cheap in money and expensive only in time. Image inputs: first_frame animates a still picture, reference_images keep a subject/style consistent — both accept a job_id of a completed image generation on this account, a public image URL, or inline base64 image data, and on omni-flash they can be combined (up to 10 images total). Example: {"prompt": "drone shot over a misty pine forest at sunrise", "model": "veo-3.1-fast", "duration": 8, "resolution": "720p", "confirm_cost": 0.70}
Schéma d’entrée
{'type': 'object', 'required': ['prompt'], 'properties': {'seed': {'type': 'integer', 'maximum': 2147483647, 'minimum': 0, 'description': 'Veo and wan-3.0 only.'}, 'model': {'enum': ['veo-3.1', 'veo-3.1-fast', 'veo-3.1-lite', 'omni-flash', 'omni-flash-1.0', 'wan-3.0', 'grok-imagine-video-1.5'], 'type': 'string', 'default': 'veo-3.1-fast'}, 'prompt': {'type': 'string', 'maxLength': 32000, 'description': 'What to film. The limit differs per model: 2000 characters on veo-3.1*, 20000 on omni-flash and wan-3.0, 2048 on grok-imagine-video-1.5. list_models reports max_prompt_chars for each.'}, 'duration': {'enum': [3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30], 'type': 'integer', 'default': 8, 'description': 'Clip length in seconds. Veo accepts only 4, 6 or 8; omni-flash accepts 3, 4, 5, 6, 7, 8, 9, 10; omni-flash-1.0 accepts 4, 6, 8, 10 (1080p: 6, 8, 10); wan-3.0 accepts 4, 6, 8, 10, 15, 20, 30; grok-imagine-video-1.5 accepts 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.'}, 'last_frame': {'type': 'string', 'description': 'Closing frame for Veo, omni-flash 1.1, and models that advertise last-frame support: the model interpolates from first_frame to this image. Requires first_frame. Same accepted forms as first_frame (job_id, public URL or base64 data URL). omni-flash-1.0 does not support it.'}, 'resolution': {'enum': ['360p', '720p', '1080p', '4k', '480p'], 'type': 'string', 'default': '720p', 'description': 'wan-3.0: 480p / 720p / 1080p; grok-imagine-video-1.5: 720p / 1080p. Veo: 720p / 1080p, plus 4k on veo-3.1 and veo-3.1-fast. omni-flash: 360p (draft, a third of the price and up to 60% faster), 720p (native), 1080p and 4k (upscaled). 360p is omni-flash only. omni-flash-1.0: 720p, or 1080p upscaled (6/8/10 s only).'}, 'with_audio': {'type': 'boolean', 'default': False, 'description': 'Native audio for Veo models (costs more). omni-flash and grok-imagine-video-1.5 always have audio. On wan-3.0 audio is on by default and free — pass false for a silent clip.'}, 'first_frame': {'type': 'string', 'description': 'Start the clip from a still image, which is then animated. Either a job_id of a completed image generation on this account (see list_generations), a public http(s) image URL, or an inline data:image/png|jpeg|webp;base64,... URL for a local file (max 10 MB; a bare local filesystem path cannot be read by this remote server). This includes a signed URL returned by get_result, which is how you pick one variant of a multi-image generation. On omni-flash it can be combined with reference_images; on Veo it cannot.'}, 'aspect_ratio': {'enum': ['16:9', '9:16', '1:1', '3:2', '2:3'], 'type': 'string', 'default': '16:9', 'description': 'Veo and omni-flash: 16:9 or 9:16. wan-3.0: 16:9 / 9:16; grok-imagine-video-1.5: 16:9 / 9:16 / 1:1 / 3:2 / 2:3.'}, 'audio_prompt': {'type': 'string', 'maxLength': 500, 'description': 'Describe the desired sound (used when audio is on).'}, 'confirm_cost': {'type': 'number', 'description': 'The quoted USD cost you accept. Omit on the first call to get the quote.'}, 'relaxed_filter': {'type': 'boolean', 'description': "Accepted for compatibility and has no effect on video: Google's video API exposes no configurable safety settings, personGeneration=allow_adult is already the default for Veo 3.1, and Omni and gateway models have no such switch at all. Video is filtered both before generation and again on the finished clip (rai_media_filtered_reasons), and neither check can be turned off. When a clip is refused, the levers that actually work are: retry (each run renders different footage, and failures are fully refunded), reword the people-related part of the prompt, or move an omni-flash idea to veo-3.1-fast, whose filter is noticeably looser."}, 'idempotency_key': {'type': 'string', 'maxLength': 64, 'description': 'Optional unique key; retries with the same key never double-charge.'}, 'negative_prompt': {'type': 'string', 'maxLength': 1000, 'description': 'What to avoid. On omni-flash it is appended to the prompt as plain text (the model has no separate negative field).'}, 'reference_images': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 10, 'description': 'Reference images that keep a subject, character or style consistent (they are NOT used as literal frames). Each item is a job_id of a completed image generation on this account, a public http(s) image URL, or an inline data:image/png|jpeg|webp;base64,... URL. For a local file, pass a base64 data URL rather than its filesystem path. omni-flash: up to 10 images in total together with first_frame and last_frame. Veo 3.1 / Fast: at most 3, only with duration 8 and without frames. Veo 3.1 Lite: unsupported. grok-imagine-video-1.5: up to 7, at 720p only and without first_frame (at 1080p use first_frame instead — references are not accepted there).'}, 'reference_videos': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 5, 'description': 'wan-3.0 only: clips the model reads before generating — to restyle or rework a clip, continue its story, or carry its characters, motion and setting into a new scene; what to do with them is said in the prompt, where they are "Video 1", "Video 2" in order (job_ids are numbered first, then URLs). Each item is a job_id of a completed video on this account, a public http(s) video URL, or a reference_video_refs value returned by the quote (reuses the already downloaded clip). Up to 5 clips and 15 s of video in total (each clip is normalised to MP4 720p and whole seconds, longer clips are trimmed). INPUT SECONDS ARE BILLED like output seconds — the provider charges for both — and video in + video out must fit in 30 s. Cannot be combined with first_frame / last_frame; reference_images are fine. For a single-clip edit or continuation edit_video with model wan-3.0 is the simpler call.'}, 'edit_from_generation_id': {'type': 'string', 'description': 'omni-flash / omni-flash-1.0 only: job_id of a completed omni video to refine conversationally; prompt describes the changes. Duration, aspect ratio, the scene and the Omni generation that shot the clip are inherited from it — duration is ignored here (use edit_video to shorten a clip).'}}, 'additionalProperties': False}
get_account
Get the current account balance (USD), this API key's name, optional daily spend cap and how much of it is used today. Free, no charge. Use it to check affordability before starting expensive generations. When the balance no longer covers even the cheapest image, the response also carries a ready-to-open top_up_url.
Schéma d’entrée
{'type': 'object', 'properties': {}, 'additionalProperties': False}
get_result
Get the status and result of a generation job started with generate_image / edit_image / generate_video / edit_video. Waits up to wait_seconds for completion before returning (long-poll). On success returns hosted media URLs (valid 24 h — call again for fresh links), cost_charged_usd and balance_remaining_usd, plus a small inline preview for images. Free, no charge. Poll roughly every 10–15 s for videos.
Schéma d’entrée
{'type': 'object', 'required': ['job_id'], 'properties': {'job_id': {'type': 'string'}, 'wait_seconds': {'type': 'integer', 'default': 15, 'maximum': 30, 'minimum': 0, 'description': 'How long to wait server-side before answering.'}}, 'additionalProperties': False}
list_generations
List this account's recent generations (both MCP and website) — id, type, model, status, cost and prompt preview. Use it to find a job_id to re-download results or to pick a source for edit_image / edit_video / generate_video edit_from_generation_id. Free, no charge.
Schéma d’entrée
{'type': 'object', 'properties': {'type': {'enum': ['image', 'video'], 'type': 'string', 'description': 'Filter by media type.'}, 'limit': {'type': 'integer', 'default': 10, 'maximum': 50, 'minimum': 1}, 'status': {'enum': ['processing', 'completed', 'failed'], 'type': 'string', 'description': 'Filter by status.'}}, 'additionalProperties': False}
list_models
List all available image, video and speech generation models with current per-unit USD prices, supported resolutions, durations and constraints. Prices come from the same source as the website — call this before quoting costs to a user or choosing a model. Free, no charge.
Schéma d’entrée
{'type': 'object', 'properties': {}, 'additionalProperties': False}
top_up
Get a balance top-up link. OAuth connections receive a one-time deposit-only link valid for 30 minutes; opening it cannot expose API keys, profile data or generation history. Bearer API-key users receive the normal profile link. Free, no charge.
Schéma d’entrée
{'type': 'object', 'properties': {}, 'additionalProperties': False}
Modifié
generate_video
1 October 2026 02:51
Modifié
edit_image
1 October 2026 02:51
Modifié
generate_image
1 October 2026 02:51
Modifié
generate_video
25 September 2026 02:59
Modifié
generate_video
23 September 2026 02:50
Modifié
edit_image
23 September 2026 02:50
Modifié
generate_image
23 September 2026 02:50
Modifié
generate_video
19 September 2026 02:48
Modifié
generate_image
19 September 2026 02:48
Ajouté
list_generations
17 September 2026 12:54
Ajouté
get_result
17 September 2026 12:54
Ajouté
generate_speech
17 September 2026 12:54
Ajouté
edit_video
17 September 2026 12:54
Ajouté
generate_video
17 September 2026 12:54
Ajouté
edit_image
17 September 2026 12:54
Ajouté
generate_image
17 September 2026 12:54
Ajouté
top_up
17 September 2026 12:54
Ajouté
get_account
17 September 2026 12:54
Ajouté
list_models
17 September 2026 12:54