BananaBanana Image, Video & Speech Generation
Was dieses MCP kann
Generates and edits images, videos, and speech using hosted AI media models, with job tracking, model selection, pricing, and account balance tools.
Tools
Eingabeschema
{'type': 'object', 'required': ['source_generation_id', 'prompt'], 'properties': {'seed': {'type': 'integer', 'maximum': 2147483647, 'minimum': 0}, 'model': {'enum': ['nano-banana-2-lite', 'nano-banana-2', 'nano-banana-pro'], 'type': 'string', 'default': 'nano-banana-2'}, 'prompt': {'type': 'string', 'maxLength': 32000, 'description': 'The edit instruction (up to 32000 characters).'}, 'resolution': {'enum': ['512', '1024', '2048', '4096'], 'type': 'string', 'default': '1024'}, 'aspect_ratio': {'enum': ['1:1', '3:2', '2:3', '4:3', '3:4', '4:5', '5:4', '9:16', '16:9', '21:9'], 'type': 'string', 'default': '1:1'}, 'output_format': {'enum': ['jpeg', 'png', 'webp'], 'type': 'string', 'default': 'jpeg'}, 'relaxed_filter': {'type': 'boolean', 'default': True, 'description': 'Enabled by default for Google images. Uses safetySettings thresholds OFF (no configurable threshold blocking) and personGeneration=ALLOW_ADULT. Pass false to use BLOCK_ONLY_HIGH and omit personGeneration. Runs only on our own Google keys: when none is free, an explicit true is refused with SERVICE_UNAVAILABLE before any charge, and the default silently falls back to the standard filter (get_result then reports relaxed_filter: false). Ignored by other image providers and by every video model. OFF and BLOCK_NONE are image/Gemini safety options, not video parameters. Non-configurable checks still apply, including checks on generated images and input media. Celebrity refusals (support codes 15236754 / 29310472) are not disabled by this flag. get_result explains the category and stage when Google supplies support codes; failed jobs are refunded.'}, 'idempotency_key': {'type': 'string', 'maxLength': 64}, 'source_generation_id': {'type': 'string', 'description': 'job_id of a completed image generation owned by this account.'}}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'required': ['prompt'], 'properties': {'mode': {'enum': ['edit', 'extend'], 'type': 'string', 'default': 'edit', 'description': '"edit" rewrites the clip; "extend" continues it. omni-flash: edit keeps the length, extend appends a 3–10 s continuation (total capped at 40 s). wan-3.0: edit rebuilds the scene at the chosen duration, extend shoots the next scene.'}, 'model': {'enum': ['omni-flash', 'omni-flash-1.0', 'wan-3.0'], 'type': 'string', 'default': 'omni-flash', 'description': 'omni-flash (Gemini Omni 1.1) keeps the motion and length of the source and can extend the scene; omni-flash-1.0 is the previous generation (edit only, same length); wan-3.0 reads the source as a reference and renders a new clip of the chosen duration (edit or extend by prompt).'}, 'prompt': {'type': 'string', 'maxLength': 32000, 'description': 'What to change in the video (up to 20000 characters).'}, 'duration': {'enum': [3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30], 'type': 'integer', 'description': 'omni-flash with mode "edit": trim the source to the first N seconds and edit only that part; values above the source length are ignored. With mode "extend": the length of the appended continuation (3–10 s). wan-3.0: the length of the OUTPUT clip — 4, 6, 8, 10, 15, 20, 30 s (default 8); source + output must fit in 30 s.'}, 'video_url': {'type': 'string', 'description': 'Public http(s) URL of the source video (mp4/mov/webm/mkv/avi/wmv/flv/3gpp, max 200 MB). Use instead of source_generation_id.'}, 'resolution': {'enum': ['480p', '720p', '1080p'], 'type': 'string', 'description': 'wan-3.0 only: output resolution (default 720p). omni-flash edits are always 720p.'}, 'source_ref': {'type': 'string', 'description': 'Returned by the quote when video_url is used: pass it back with confirm_cost to reuse the already downloaded clip instead of downloading it again.'}, 'with_audio': {'type': 'boolean', 'default': True, 'description': 'wan-3.0 only: sound is on by default and free — pass false for a silent clip. omni-flash always generates audio.'}, 'audio_prompt': {'type': 'string', 'maxLength': 500, 'description': 'Describe the desired sound — Omni Flash always generates audio.'}, 'confirm_cost': {'type': 'number', 'description': 'The quoted USD cost you accept. Omit on the first call to get the quote.'}, 'idempotency_key': {'type': 'string', 'maxLength': 64, 'description': 'Optional unique key; retries with the same key never double-charge.'}, 'reference_images': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 10, 'description': 'omni-flash with mode "extend" only: images of subjects, products or characters the continuation should bring into the scene — a completed image job_id, a public http(s) URL, or a data:image/...;base64,... value (≤10 MB). Refer to them in the prompt ("the item from the reference image enters the frame"). Ignored by mode "edit" and by wan-3.0.'}, 'source_generation_id': {'type': 'string', 'description': 'job_id of a completed video generation on this account.'}}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'required': ['prompt'], 'properties': {'seed': {'type': 'integer', 'maximum': 2147483647, 'minimum': 0, 'description': 'For reproducible results.'}, 'model': {'enum': ['nano-banana-2-lite', 'nano-banana-2', 'nano-banana-pro', 'gpt-image-2.5-flare', 'gpt-image-2.5-sunburst', 'qwen-image-3.0-pro'], 'type': 'string', 'default': 'nano-banana-2-lite', 'description': 'nano-banana-2-lite: cheapest Google default, 1024 only. nano-banana-2: choose for 512, 2048 or 4096 output. nano-banana-pro: top Google quality, up to 4K. gpt-image-2.5-flare / gpt-image-2.5-sunburst: OpenAI — strongest at readable in-image text and long literal briefs; 1024, 2048 or 4096 (4K = 3840 on the long side); Flare is the fast 2.5 tier, Sunburst the precision tier; they ignore seed and relaxed_filter, and with reference_images they take the orientation of the first reference (omit aspect_ratio there, it is rejected). qwen-image-3.0-pro: Alibaba — crisp small text, dense layouts (posters, menus, UI); 1024 or 2048 only, up to 3 references, seed + negative_prompt. Cheapest per image: nano-banana-2-lite at $0.03.'}, 'prompt': {'type': 'string', 'maxLength': 32000, 'description': 'What to generate. English works best. Length limit depends on the model — 32000 characters on every image model here; list_models reports max_prompt_chars for each.'}, 'resolution': {'enum': ['512', '1024', '2048', '4096'], 'type': 'string', 'default': '1024', 'description': 'Lite requires 1024. Choose nano-banana-2 for 512, 2048 or 4096; 4096 costs the most. GPT Image models have no 512; qwen-image-3.0-pro accepts 1024 or 2048 only.'}, 'aspect_ratio': {'enum': ['1:1', '3:2', '2:3', '4:3', '3:4', '4:5', '5:4', '9:16', '16:9', '21:9'], 'type': 'string', 'default': '1:1'}, 'confirm_cost': {'type': 'number', 'description': 'Required for batches (number_of_images > 1): the quoted total USD cost you accept.'}, 'output_format': {'enum': ['jpeg', 'png', 'webp'], 'type': 'string', 'default': 'jpeg'}, 'relaxed_filter': {'type': 'boolean', 'default': True, 'description': 'Enabled by default for Google images. Uses safetySettings thresholds OFF (no configurable threshold blocking) and personGeneration=ALLOW_ADULT. Pass false to use BLOCK_ONLY_HIGH and omit personGeneration. Runs only on our own Google keys: when none is free, an explicit true is refused with SERVICE_UNAVAILABLE before any charge, and the default silently falls back to the standard filter (get_result then reports relaxed_filter: false). Ignored by other image providers and by every video model. OFF and BLOCK_NONE are image/Gemini safety options, not video parameters. Non-configurable checks still apply, including checks on generated images and input media. Celebrity refusals (support codes 15236754 / 29310472) are not disabled by this flag. get_result explains the category and stage when Google supplies support codes; failed jobs are refunded.'}, 'idempotency_key': {'type': 'string', 'maxLength': 64, 'description': 'Optional unique key; retries with the same key never double-charge.'}, 'negative_prompt': {'type': 'string', 'maxLength': 1000}, 'number_of_images': {'type': 'integer', 'default': 1, 'maximum': 4, 'minimum': 1, 'description': 'Variants per call. >1 requires confirm_cost.'}, 'reference_images': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 14, 'description': "Actual visual references for the generated image. Each item is a job_id of a completed image generation on this account, a public http(s) image URL, or an inline data:image/png|jpeg|webp;base64,... URL (max 10 MB each). For a local file, read and base64-encode its bytes into a data URL; a bare local path such as C:\\\\photo.jpg cannot be read by this remote server. Google and Qwen models keep aspect_ratio with references; GPT Image models follow the first reference's orientation instead and reject an explicit aspect_ratio."}}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'required': ['text'], 'properties': {'text': {'type': 'string', 'maxLength': 6000, 'description': 'Exact transcript to speak. For dialogue, prefix every turn with the matching speaker name, e.g. Sam: Hello.'}, 'style': {'type': 'string', 'maxLength': 1500, 'description': 'Optional overall direction: persona, scene, emotion, accent, pace, pronunciation and delivery notes.'}, 'voice': {'enum': ['Achernar', 'Achird', 'Algenib', 'Algieba', 'Alnilam', 'Aoede', 'Autonoe', 'Callirrhoe', 'Charon', 'Despina', 'Enceladus', 'Erinome', 'Fenrir', 'Gacrux', 'Iapetus', 'Kore', 'Laomedeia', 'Leda', 'Orus', 'Pulcherrima', 'Puck', 'Rasalgethi', 'Sadachbia', 'Sadaltager', 'Schedar', 'Sulafat', 'Umbriel', 'Vindemiatrix', 'Zephyr', 'Zubenelgenubi'], 'type': 'string', 'default': 'Kore', 'description': 'Single-speaker voice. Ignored when speakers is provided.'}, 'speakers': {'type': 'array', 'items': {'type': 'object', 'required': ['name', 'voice'], 'properties': {'name': {'type': 'string', 'maxLength': 40, 'minLength': 1}, 'voice': {'enum': ['Achernar', 'Achird', 'Algenib', 'Algieba', 'Alnilam', 'Aoede', 'Autonoe', 'Callirrhoe', 'Charon', 'Despina', 'Enceladus', 'Erinome', 'Fenrir', 'Gacrux', 'Iapetus', 'Kore', 'Laomedeia', 'Leda', 'Orus', 'Pulcherrima', 'Puck', 'Rasalgethi', 'Sadachbia', 'Sadaltager', 'Schedar', 'Sulafat', 'Umbriel', 'Vindemiatrix', 'Zephyr', 'Zubenelgenubi'], 'type': 'string'}}, 'additionalProperties': False}, 'maxItems': 2, 'minItems': 2, 'description': 'Exactly two dialogue speakers. Their names must prefix the turns in text.'}, 'language_code': {'type': 'string', 'description': 'Optional language/locale such as en-US, ru-RU, ja-JP or es-MX. Omit for automatic detection.'}, 'idempotency_key': {'type': 'string', 'maxLength': 64, 'description': 'Optional unique key; retries with the same key never double-charge.'}}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'required': ['prompt'], 'properties': {'seed': {'type': 'integer', 'maximum': 2147483647, 'minimum': 0, 'description': 'Veo and wan-3.0 only.'}, 'model': {'enum': ['veo-3.1', 'veo-3.1-fast', 'veo-3.1-lite', 'omni-flash', 'omni-flash-1.0', 'wan-3.0', 'grok-imagine-video-1.5'], 'type': 'string', 'default': 'veo-3.1-fast'}, 'prompt': {'type': 'string', 'maxLength': 32000, 'description': 'What to film. The limit differs per model: 2000 characters on veo-3.1*, 20000 on omni-flash and wan-3.0, 2048 on grok-imagine-video-1.5. list_models reports max_prompt_chars for each.'}, 'duration': {'enum': [3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 30], 'type': 'integer', 'default': 8, 'description': 'Clip length in seconds. Veo accepts only 4, 6 or 8; omni-flash accepts 3, 4, 5, 6, 7, 8, 9, 10; omni-flash-1.0 accepts 4, 6, 8, 10 (1080p: 6, 8, 10); wan-3.0 accepts 4, 6, 8, 10, 15, 20, 30; grok-imagine-video-1.5 accepts 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.'}, 'last_frame': {'type': 'string', 'description': 'Closing frame for Veo, omni-flash 1.1, and models that advertise last-frame support: the model interpolates from first_frame to this image. Requires first_frame. Same accepted forms as first_frame (job_id, public URL or base64 data URL). omni-flash-1.0 does not support it.'}, 'resolution': {'enum': ['360p', '720p', '1080p', '4k', '480p'], 'type': 'string', 'default': '720p', 'description': 'wan-3.0: 480p / 720p / 1080p; grok-imagine-video-1.5: 720p / 1080p. Veo: 720p / 1080p, plus 4k on veo-3.1 and veo-3.1-fast. omni-flash: 360p (draft, a third of the price and up to 60% faster), 720p (native), 1080p and 4k (upscaled). 360p is omni-flash only. omni-flash-1.0: 720p, or 1080p upscaled (6/8/10 s only).'}, 'with_audio': {'type': 'boolean', 'default': False, 'description': 'Native audio for Veo models (costs more). omni-flash and grok-imagine-video-1.5 always have audio. On wan-3.0 audio is on by default and free — pass false for a silent clip.'}, 'first_frame': {'type': 'string', 'description': 'Start the clip from a still image, which is then animated. Either a job_id of a completed image generation on this account (see list_generations), a public http(s) image URL, or an inline data:image/png|jpeg|webp;base64,... URL for a local file (max 10 MB; a bare local filesystem path cannot be read by this remote server). This includes a signed URL returned by get_result, which is how you pick one variant of a multi-image generation. On omni-flash it can be combined with reference_images; on Veo it cannot.'}, 'aspect_ratio': {'enum': ['16:9', '9:16', '1:1', '3:2', '2:3'], 'type': 'string', 'default': '16:9', 'description': 'Veo and omni-flash: 16:9 or 9:16. wan-3.0: 16:9 / 9:16; grok-imagine-video-1.5: 16:9 / 9:16 / 1:1 / 3:2 / 2:3.'}, 'audio_prompt': {'type': 'string', 'maxLength': 500, 'description': 'Describe the desired sound (used when audio is on).'}, 'confirm_cost': {'type': 'number', 'description': 'The quoted USD cost you accept. Omit on the first call to get the quote.'}, 'relaxed_filter': {'type': 'boolean', 'description': "Accepted for compatibility and has no effect on video: Google's video API exposes no configurable safety settings, personGeneration=allow_adult is already the default for Veo 3.1, and Omni and gateway models have no such switch at all. Video is filtered both before generation and again on the finished clip (rai_media_filtered_reasons), and neither check can be turned off. When a clip is refused, the levers that actually work are: retry (each run renders different footage, and failures are fully refunded), reword the people-related part of the prompt, or move an omni-flash idea to veo-3.1-fast, whose filter is noticeably looser."}, 'idempotency_key': {'type': 'string', 'maxLength': 64, 'description': 'Optional unique key; retries with the same key never double-charge.'}, 'negative_prompt': {'type': 'string', 'maxLength': 1000, 'description': 'What to avoid. On omni-flash it is appended to the prompt as plain text (the model has no separate negative field).'}, 'reference_images': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 10, 'description': 'Reference images that keep a subject, character or style consistent (they are NOT used as literal frames). Each item is a job_id of a completed image generation on this account, a public http(s) image URL, or an inline data:image/png|jpeg|webp;base64,... URL. For a local file, pass a base64 data URL rather than its filesystem path. omni-flash: up to 10 images in total together with first_frame and last_frame. Veo 3.1 / Fast: at most 3, only with duration 8 and without frames. Veo 3.1 Lite: unsupported. grok-imagine-video-1.5: up to 7, at 720p only and without first_frame (at 1080p use first_frame instead — references are not accepted there).'}, 'reference_videos': {'type': 'array', 'items': {'type': 'string'}, 'maxItems': 5, 'description': 'wan-3.0 only: clips the model reads before generating — to restyle or rework a clip, continue its story, or carry its characters, motion and setting into a new scene; what to do with them is said in the prompt, where they are "Video 1", "Video 2" in order (job_ids are numbered first, then URLs). Each item is a job_id of a completed video on this account, a public http(s) video URL, or a reference_video_refs value returned by the quote (reuses the already downloaded clip). Up to 5 clips and 15 s of video in total (each clip is normalised to MP4 720p and whole seconds, longer clips are trimmed). INPUT SECONDS ARE BILLED like output seconds — the provider charges for both — and video in + video out must fit in 30 s. Cannot be combined with first_frame / last_frame; reference_images are fine. For a single-clip edit or continuation edit_video with model wan-3.0 is the simpler call.'}, 'edit_from_generation_id': {'type': 'string', 'description': 'omni-flash / omni-flash-1.0 only: job_id of a completed omni video to refine conversationally; prompt describes the changes. Duration, aspect ratio, the scene and the Omni generation that shot the clip are inherited from it — duration is ignored here (use edit_video to shorten a clip).'}}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'required': ['job_id'], 'properties': {'job_id': {'type': 'string'}, 'wait_seconds': {'type': 'integer', 'default': 15, 'maximum': 30, 'minimum': 0, 'description': 'How long to wait server-side before answering.'}}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'properties': {'type': {'enum': ['image', 'video'], 'type': 'string', 'description': 'Filter by media type.'}, 'limit': {'type': 'integer', 'default': 10, 'maximum': 50, 'minimum': 1}, 'status': {'enum': ['processing', 'completed', 'failed'], 'type': 'string', 'description': 'Filter by status.'}}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
Eingabeschema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
Letzte Tool-Änderungen
Ähnliche MCP-Server
MiOffice — AI-Powered Workspace Studio
Provides browser-based tools for processing PDFs, images, video, and audio, including generation, enhancement, conversion, transc…
BlitzReels Video Editor
Creates and edits short-form videos, including timelines, captions, transitions, AI-generated visuals, music, voiceovers, sound e…
Morpha
Enables creating, animating, and exporting layered short-form video projects.
switch
Enables management and exploration of an account-scoped library of AI-generated images and videos.
framesail
Creates long-form YouTube videos through script generation, storyboarding, visual assets, voiceover, music, scene composition, an…
Magnific
Provides image enhancement and design tools, AI image and audio generation, text-to-speech, creation management, and reusable cre…
Uwear
Supports AI fashion photoshoot production with garments, outfits, avatars, locations, art direction, image generation, backdrops,…
Creative Claw
Generates and edits branded images, video, audio, speech, and 3D-related creative assets, with media processing and reusable them…