Skip to main content

Studio Prompting Guide

Abacus AI Studio turns a written prompt, and optionally a few reference files, into images and videos. Modern image and video generative models each have their own limitations, but you can improve how well the given model fulfills a request with proper prompting: what it describes, what it leaves out, and what it tells the model to keep. This guide explains how Abacus AI Studio processes a prompt, how to structure prompts for today's image and video models, how to work from reference images, how to get results you can reproduce, and what to do when a result goes wrong.


How Studio Processes Your Prompt​

Auto chooses a model for you​

For most everyday image and video tasks, it is recommended you use the Auto mode. Auto automatically routes your request to the best suited model for your request, balanced for quality and cost. The model chosen is driven by what you ask for: whether you are creating from scratch or editing something you attached, whether the result needs legible text, a real-world subject, or a transparent background, and for video, whether you supplied an opening frame, reference images, or a clip to change. You can adjust the tradeoff between quality and cost by specifying in your prompt the desired quality: Standard, Medium, High, or Max.

Because the routing choice is driven by your wording, describe plainly what you want ("edit the attached image", "transparent background", "animate this character", "use this image as a start frame", etc.) rather than hinting at it. If you want a specific model, pick it in the model dropdown, or name it explicitly in the prompt. Picking your own model is the right choice when you want the most consistency and repeatability across many successive generations.

Rewrite Prompt​

Under the "…" settings there is a Rewrite Prompt switch. When ON, the Abacus AI Studio agent expands your prompt into a fuller description before it reaches the image or video model. The agent is tuned to properly add detail specific to model chosen to maximize the quality of the results, however it may invent details not present in the original prompt for underspecified requests. The Rewrite Prompt switch is best used for short prompts where output quality is more important than specific details in the prompt. For a detailed, carefully written prompt where specifics matter, it's best to turn it off, as results may invent new details and differ slightly between generations.

  • Any request with an attached image (edits, references, start frames): the rewrite is skipped automatically, whatever the switch says. Your words go straight to the model to maximize accuracy.
  • MiniMax H3 additionally has a Provider Prompt Optimizer that rewrites on MiniMax's side. Set it to Disabled when exact wording matters.
  • Midjourney prompts are always rewritten, and only these inline flags survive: --sref <code|random>, --sw, --seed, --stylize, --chaos, --weird, --tile, --style raw, --no. Aspect ratio, version and reference URLs are stripped; set those in the pills.

Attachments​

Studio has the ability to include reference files in your requests: add files from your computer or from previous Studio conversations. Studio works out each attachment's role from your prompt: "edit the attached image" makes it the thing being edited, "the product in the attached image, placed on …" makes it a subject reference, "in the style of the attached image" makes it a style reference, and for video "use the attached image as the first frame" versus "the character from the attached image walks into …" decides start frame versus identity reference.

Attachments are numbered in the order you add them, and both the Studio agent as well as the generative model understand this order. Use "image 1", "image 2", "video 1", etc. to refer to specific files in your prompt. Add them in the order you will refer to them.

Aspect ratio and resolution​

  • Editing an attached reference: leaving Aspect Ratio on Auto will result in the output matching the source. Setting a different ratio forces a re-composition.
  • From scratch: set the ratio explicitly through the dropdown. Using Auto here will result in the model choosing a ratio based on the request.
  • Resolution: Higher resolutions generally consume more credits. In extreme cases, ultra-high resolution requests may result in less-capable models being used, as not all premium models can generate at 4K and higher. It is recommended to request a resolution only as hight as necessary for your use case to allow flexibility to use the best models available.

Writing an Image Prompt​

Structure​

Write the prompt the way a photographer would brief a shoot: in sentences rather than keywords, with the most important thing first. Models weight the start of a prompt more than the end, and they read full sentences far better than a comma-separated pile of tags. A good prompt names, roughly in this order:

  • Format and intent: "editorial product photograph", "architectural photograph", "flat vector icon". This sets the level of polish and stops the model guessing.
  • Subject with concrete, material detail: "a matte black ceramic pour-over kettle with a walnut handle", not "a nice kettle".
  • Composition, lighting and camera: shot type and angle, where the light comes from and how soft it is, focal length and aperture.
  • Style and finish: colour grading, mood, era.
  • Text, in double quotes, with its placement and typeface style.
  • Constraints, phrased positively where possible.

Thirty to eighty words is the sweet spot for a from-scratch image. Once prompts become too long, image models begin to average the instructions instead of following all of them faithfully. Here is an example:

Editorial product photograph of a matte black ceramic pour-over kettle with a walnut handle on a pale oak counter. Three-quarter view at eye level, kettle slightly left of centre, a single white cup out of focus behind it. Soft north-facing window light from the left, gentle shadow to the right. 50mm lens at f/2.8, shallow depth of field. Natural colour, no colour grading, plain background with nothing else on the counter.

Realism​

Models produce a photograph when the prompt reads like the description of one. Borrow the vocabulary of real photography rather than asking for "realistic":

  • Camera and lens: "35mm lens at f/8", "85mm portrait lens, shallow depth of field", "24mm wide angle at eye level, slight barrel distortion".
  • Light with a direction: "soft north window light from the left", "late-afternoon sun from the west, long shadows", "single softbox from upper left, no fill".
  • Texture and imperfection: "visible skin pores, unretouched", "fabric wear", "dust and fingerprints on the glass", "subtle weathering", "shot on 35mm film, fine grain".

On GPT Image 2 the word "photorealistic" itself switches the model into its photographic mode. Avoid quality adjectives such as "8K", "ultra-detailed", "masterpiece", "stunning" or "flawless": they carry no visual information and push results toward an over-sharpened, over-saturated, plastic render. "Perfect skin" and "flawless" in particular tend to produce outputs that look less realistic.

Text and exclusions​

Put the exact words in double quotes and say where they go and what the type looks like: the headline "SPRING SALE" top centre in a heavy geometric sans-serif, white on the green band. Single words and short phrases are reliable; paragraphs are not. For a logo or wordmark that already exists, attach it rather than describing it (see Working From References).

Phrase exclusions positively where you can: "an empty street" rather than "no cars", "a plain seamless background" rather than "no clutter". GPT Image 2 also honours a plain "no watermark, no extra text, no people"; Midjourney takes --no; Nano Banana, Seedream and FLUX have no negative prompt at all and respond only to the positive form.


Writing a Video Prompt​

Structure​

Subject + action + scene + camera + lighting, 80 to 150 words, one continuous shot. Front-load the subject and the movement.

  • Specify the camera. "Slow dolly in", "static locked-off camera", "tracking shot from the side at knee height". Without it, models invent random framing.
  • Give motion an endpoint. "Turns her head to the left, then holds" rather than "turns her head". Open-ended motion drifts.
  • One camera move and one main action per clip. Stacked actions and stacked camera moves are what cause morphing and warped limbs. Three locations in a five-second clip becomes smeared morphing.
  • Gradual beats sudden. Fast stunts, jumps and whip cuts produce artifacts.
  • Avoid on-screen text. Video models tend to render text poorly. If a sign or title must be legible, generate a still with that text in an image model first and include it as the start frame or reference.
  • Name the audio or the model invents it. "Ambient street sound, no music, no dialogue, no subtitles." Video models may invent random speech, music, or background sounds unless specified not to. Additionally specify language for all speaking video requests.
  • Real people and specific products need a reference image. A name alone will not render a recognisable person.

A text-to-video prompt in full:

A barista in a small sunlit café pours steamed milk into a ceramic cup, finishing a rosetta as the pour ends. Static camera, locked off, medium close-up across the counter at eye level. Soft morning light from the window on the left, warm reflections on the copper machine behind. Shallow depth of field, natural colour. Ambient café sound only, no music, no dialogue, no subtitles.

Quality tiers​

Video generation can consume a large amount of credits. The Standard and Medium tiers do the best job of balancing cost and quality, and it is recommended to stay on these tiers unless the shot's make-or-break element requires advanced physics interactions or very realistic and close up human faces or lifelike speaking. Ask for 1080p or 4K only when the deliverable needs it; both re-price the clip. Higher quality tiers can also handle more complex prompts, and High or Max settings can allow for multiple beats or shots per clip, with realistic speaking and character consistency.


Audio Prompting​

  • Music: describe the track itself rather than the scene it accompanies. Specify genre, mood, instrumentation, tempo and, for songs, the vocal style ("warm lo-fi hip hop, mellow, Rhodes piano and soft brushed drums, slow tempo, instrumental"). For sung vocals, supply the lyrics one line per line, using [Verse], [Chorus] and [Bridge] tags; the length of the song follows the lyrics. For a background bed under narration, state "instrumental" explicitly, since the model may otherwise add vocals. Track length is approximate on the default model; select a model with exact duration control when the length must be precise.
  • Voiceover: provide the text exactly as it should be read and select a voice from the voice picker. Punctuation controls pacing: commas and full stops add pauses, a dash adds a beat, and a paragraph break adds a longer pause. Write numbers, abbreviations and unusual names as they should be pronounced. The more expressive voices accept inline delivery tags such as [whispers] or (laughs). Scripts are limited to a few thousand characters per generation, so long narration should be split into sections.
  • Ambience and sound effects: describe a single non-musical sound or soundscape ("dense forest with birdsong and a distant stream", "single wooden door creak"). Ambience beds can run for several minutes; individual effects are short, and leaving the duration unset lets the model choose the natural length.

Working From References​

Studio has the ability to work from reference files in your requests -- you can add files from your computer, from previous Studio conversations, or reference previous generations from the same conversation in your prompt. Studio works out each attachment's role from your prompt: "edit the attached image" makes it the thing being edited, "the product in the attached image, placed on …" makes it a subject reference, "in the style of the attached image" makes it a style reference, and for video "use the attached image as the first frame" versus "the character from the attached image walks into …" decides start frame versus identity reference. If using a reference from the same conversation, be very specific about which reference you are referring to, or optionally attach it in the chat explicitly.

Editing a reference​

Many edits tend to include unwanted edits included in the regeneration. To avoid excess changes, explicitly state what to keep the same in the re-render. Restate this lock list on every follow-up turn. Each edit re-renders the whole image, and the model may not remember constraints from earlier turns; the lock list is what keeps the re-render honest. For the best adherence to instructions, change only one thing per turn. Bundled edits ("recolour it, move it left and swap the background") tend to cause more drift or imperfect instruction following.

Change the sneaker's upper to forest green suede. Keep the exact silhouette, sole, laces, logo placement, camera angle, lighting and the seamless pale grey background from the attached image. Change nothing else.

Deterministic edits​

Several common edits can be performed deterministically by the Abacus AI Studio agent, and do not require a generation from an image or video model. These result tend to be more accurate to the request, as input preservation and behavior is much more precisely defined. Requesting them plainly routes them to this path.

  • Images: crop, resize, rotate or flip, pad the canvas to a new aspect ratio, convert the file format, and place a logo, watermark or line of text at an exact position on an otherwise unchanged image ("place the logo from image 1 in the bottom-right corner at 10% width; keep everything else pixel-identical").
  • Video: trim, split, reorder and combine clips, change speed between 0.25x and 4x, crop or punch in, change the canvas shape, add or restyle text and captions, add transitions, fades and volume changes, and place music under a clip. These are performed in the video editor, either by requesting them in the chat with the clip attached ("trim the first two seconds and add a 0.5s cross-fade") or by opening the clip in the editor pane and adjusting the timeline directly. Credits are charged only on export.

Upscaling and vectorising are not deterministic: upscaling adds detail with a model (Magnific for a faithful upscale, Topaz for video), and converting a raster to SVG runs a vectoriser. Wording determines the route: "trim", "crop", "resize", "add the logo" and "make it editable" are handled deterministically; "restyle", "change the setting", "swap the subject" and "animate" are sent to a model. Combining deterministic and non-deterministic edits into a single prompt will likely result in a singular regeneration with all changes applied. For the best results, split the generative edits (where the image/video content is changing) from the deterministic edits (crop, resize, captions) into separate prompts.

Several references​

It is recommended not to use more than 4 references unless truly necessary; more references can dilute each other and lead to poorer adherence to each additional piece of reference material. Give every reference one job, by number and content, and put the must-preserve item first, as earlier references tend to be honored better by the generative model than later ones. An example:

Image 1 is the product and must be preserved exactly. Image 2 is the lighting and colour reference only. Image 3 is the environment. Place the product from image 1 on the surface in image 3, lit as in image 2, matching shadows and colour temperature so it does not look pasted on.

Video models take the same instruction, with their own citation form: @Image1 on Seedance and Kling, "Image 1" on MiniMax, always in attachment order. It is advised to not let multiple references control the same thing: a daylight product shot plus a night-time mood board can produce a muddy compromise. Additionally, reference quality matters more than count. A clean, well-lit, uncropped file of at least 1024px on the long edge beats multiple weaker ones.

Video: start frames, reference images and elements​

A video attachment plays one of three roles, and the prompt decides which.

  • Start frame ("Use as Start Frame", or "use the attached image as the first frame"): the clip opens on exactly this image. The prompt should describe only what moves and what changes, never re-describe the picture; re-describing the frame fights it and causes drift. Leave aspect ratio unset so it follows the frame.
  • Reference images: the subject appears in a new scene. Name each reference's role and what to preserve: "the woman in image 1, keeping her face, hair and jacket exactly, walks through the office in image 2". On Kling O3, Elements (under "…") go a step further: one frontal photo plus one to three more angles per character or product, cited as @Element1, holds identity across shots and camera moves.
  • Start and end frame: the clip interpolates from one still to the other. Keep the two frames compositionally close; big jumps morph.

Frames and reference images cannot be combined on most models, so pick one approach per clip. Seedance rejects photorealistic human faces in uploaded references; use Kling O3 for real people, or a stylised reference.

Consistency across a series​

To keep a character, product or set consistent across a campaign, make one hero image you are happy with, then generate every variation as an edit of that hero ("Use as reference image") rather than from a fresh text prompt. Name the subject the same way each time ("the K2 kettle") with a one-line description, and attach three to five angles of the subject when the pose must change. If drift creeps in, go back to the hero and restart from it rather than editing the drifted image further. Video works the same way: reuse the same reference images or elements across every clip, and to continue a scene attach the finished clip and ask Studio to continue it so the next clip starts from its last frame. For the best consistency, set an explicit image or video model, and use the same model for all generations. If you are using a model with a Seed parameter, setting that will also improve consistency across generations.

Logos, wordmarks and exact text should never be described: attach the file, refer to it as "the logo from image 1", describe only its placement and size, and composite the real asset into final deliverables. Where the Edit Image brush is available, mask the region to change and describe only that region; without a mask, "edit only the background" in the scope line does most of the same work.


Shorts and Avatars​

Shorts​

A Short is a vertical, single-take presenter video built from an avatar and, optionally, a template and/or a product. The text input is a brief, not a script: Studio researches the product, writes the concept, beats and spoken lines, and presents the plan for approval before rendering. A brief is most effective when it is short and specific about what the format cannot infer: the setting, the single feature or claim to lead with, the tone, and anything to avoid. For the best and most repeatable results, it is recommended to select one of the premade templates, which excel at showcasing specific common video formats. When a chosen template already covers the intent, the brief can be left empty.

To have specific words spoken, include the actual lines. Studio detects a written script and reproduces it verbatim in place of its own, so lines should be provided beat by beat rather than as a topic. Keep them short: a 15-second Short carries roughly two or three sentences of speech, and a 30-second Premium Short about twice that. On-screen text, captions and music are determined by the format and should not be requested in the brief. The spoken language follows the language of the script.

Selfie Review, in a bright kitchen in the morning. Lead with how quiet the kettle is. Casual, slightly amused, no hard sell.

Avatars​

An avatar is a reusable presenter, a face and optionally a voice, that Shorts uses as its reference for every video made with it. Reusing the same avatar is what keeps the presenter consistent across a campaign. An avatar is created either from a description ("warm late-20s creator, shoulder-length dark hair, casual denim jacket, friendly energy") or from up to three uploaded photos of one person. Uploads should be clear, well-lit, front-facing portraits with the face fully visible; group shots, sunglasses and heavy filters weaken the likeness. Avatars cannot depict real people or public figures, so the description must be of a fictional person. The voice is created either from a description of how it sounds ("bright American accent, upbeat and conversational, talking straight to camera") or from a clean 30 to 60 second recording of a single speaker with no background music.


Repeatability Playbook​

What actually repeats​

The most-used image models have no seed: GPT Image 2 and the Nano Banana family return a fresh sample on every run, however carefully you copy the prompt. Seedream 5 and FLUX 2 Pro expose a Seed field, and Midjourney takes --seed inline, but even there the same seed gives a similar composition rather than an identical image, and only while every other setting stays the same. Video models in Studio expose no seed either. Repeatability therefore primarily comes from inputs:

  • reuse the same model for all generations
  • use a locked reference image, ideally an approved output attached as the reference for the next edit;
  • use a verbatim prompt, with Rewrite Prompt off;
  • pinned settings: aspect ratio, resolution, quality, and the order of attachments (for video, duration and the start frame as well).
  • optionally keep a shared style block and change only the subject line between generations.

When a result is right, use Copy on the result card to capture the exact prompt, and Reuse prompt or Recreate to reload the prompt, settings, model and attachments into the composer in one click.

Iterating and brand colours​

When a result is close, change exactly one thing: the lighting line, the lens, the background, the reference order. Changing multiple aspects of a generation with a single edit can lead to drift, as generative models may aggregate the request and over or under modify the input.

Models can get close to a hex code, not exact, and even the best current models score only "close enough" on brand-colour benchmarks. To maximise adherence, attach a flat swatch image as a reference, name the colour in words as well as by code ("deep cobalt blue, #1F3A93"), and bind it to a specific object ("the packaging band is #1F3A93") rather than asking for the colour "somewhere". Treat exact colour as a post-production step for final assets.


Image Model Cheat Sheet​

GPT Image 2. Natural-language paragraphs; order scene, subject, details, constraints. Best text rendering and instruction following; the strongest generalist editor, but it re-renders every pixel, so lock lists matter. Up to 16 reference images. Include "photorealistic" for photo mode. Quality: Medium for most, High for small text, faces and identity-sensitive edits. Only model family with a real transparent background. No seed. Guide: OpenAI image prompting guide.

Nano Banana Pro. Excellent at preserving what is in the source while changing what you name; the natural choice for render-to-photo, material swaps and multi-image composites. Up to 14 references; character consistency works best with 4 to 5 photos of the subject. Positive phrasing only, no negatives. Web search on by default for real-world subjects. No seed. Small text and precise diagrams are its weak spots. Guides: Google's Nano Banana Pro prompting tips, Gemini image generation docs.

Nano Banana 2. Same prompt grammar as Pro, cheaper and faster, native 4K, up to 14 references, good at typography, infographics and marketing layouts. No transparent output. Use for "enhance to 2K/4K, preserve the original exactly" upscales where slight re-rendering is acceptable. Guide: Gemini image generation docs.

Seedream 5 Pro / Lite. Reasoning model: full sentences, each clause a directive, under about 600 words. Strong at region-scoped edits ("Using the material from image 1 and the colour from image 2, modify the sofa in image 3. Keep everything else the exact same."), text in 14 languages. Has a seed field. Refer to references by content ("the tan leather journal") as well as number. Guides: BytePlus Seedream 5.0 Pro tutorial, BytePlus Seedream prompt guide (written for 4.x; the same grammar applies).

FLUX 2 Pro. Cheapest photorealism for people and products; word order matters (main subject first), 30 to 80 words ideal. Honours hex codes bound to objects. Up to 9 references with explicit roles. Has a seed. No negatives. Guide: Black Forest Labs FLUX.2 prompting guide.

Midjourney. Aesthetic exploration, not precision: literal instruction following and exact text are weaker. Prompt: subject, details, context, style, technical, 50 to 150 words; --style raw and --stylize 0-50 for photographic looks. Style codes via --sref, exclusions via --no. Image URLs, --cref and --ar are stripped in Studio; use the aspect ratio pill. Guide: Midjourney Prompt Basics.

Grok Imagine Image 2. Fast, cheap edits and mockups; write a short design brief; up to 3 references, extras are dropped. Guide: xAI image generation docs.


Video Model Cheat Sheet​

Seedance 2.0. Ignores timestamps; pace multi-beat clips with "Shot 1: … Shot 2: …". References are cited as @Image1 / @Video1 / @Audio1 in the order attached, with a role for each ("@Image1 is the protagonist; her voice is @Audio1"). Up to 9 images, 3 videos and 3 audio clips. Single-view subject photos work best. Content filter: rejects photorealistic human faces in uploaded references, named celebrities, brands and violent keywords. Cinematic or illustrated subjects pass. Guides: BytePlus Seedance 2.0 tutorial, fal Seedance 2.0 prompting guide.

Seedance 2.5 (the highest-quality single take, up to 30 seconds). The one-shot rule does not apply: write a structured brief. One summary sentence (subject, location, event, style, camera), then a gapless timeline ("0-3s: … 3-8s: …", one beat per 2 to 4 seconds), then closing notes for camera, lighting and audio that hold throughout. Bind every reference by slot and role; an uncited reference causes character confusion. Timestamps are budgets, not frame-accurate cuts; under-filled ranges improvise and over-packed ranges cause cuts. Edits: "Replace/Add/Remove … in @Video1, from 4-6s, leave the rest unchanged". Extensions: "Extend @Video1 forward: …". Negatives work only for subtitles and audio ("No subtitles", "No BGM"). Guides: BytePlus Seedance 2.5 prompt guide, fal Seedance 2.5 prompting guide.

Kling O3. Think in shots: "Shot 1 (3s): wide shot, … Shot 2 (2s): close-up, …", up to 15 seconds. Cite references as @Image1 and elements as @Element1. Elements (under "…") are the strongest consistency tool: one frontal photo plus 1 to 3 more angles per character or product, up to 7 per clip; reference them by @Element1 and the identity holds across shots and camera moves. Dialogue: @Element1 says, "…" with tone and language in parentheses. For video-to-video say "maintain the original camera movement and timing" or the model treats them as negotiable. Set the shot mode to Customize when you want your shot list followed exactly. Guides: Kling 3.0 Omni user guide, Kling 3.0 user guide.

MiniMax H3 / H3 Max. Strong prompt adherence with structured prompts; it reads the whole prompt as a language model. Cite references as "Image 1", "Video 1", "Audio 1" (up to 9 / 3 / 3). Camera as natural sentences: "The camera pushes in with small amplitude at slow speed toward the kettle." Describe the audio landscape in one to four sentences and name the music by instrumentation, not mood. The 2K default upscales a 768p render; small faces lose detail, so frame closer instead of prompting longer. Turn Provider Prompt Optimizer off for exact wording. Guide: MiniMax H3 prompting reference.

Gemini Omni Flash 1.1. Keep edit prompts short: "Change the sign to say 'OPEN'. Keep everything else the same." Over-describing causes unintended changes. Say "one continuous shot, no cuts" or it adds cuts. Source clips up to 10 seconds; voice cannot be edited. Guides: Gemini Omni prompt guide, Gemini Omni docs.


Troubleshooting​

  • The output ignores part of the prompt: it is too long. Cut to under 80 words for an image or 150 for a clip, or split the request across two turns.
  • An edit changed things you did not ask about: add a lock list, scope the edit, set Aspect Ratio to Auto, and restate the locks on every turn.
  • The same prompt looks different every run: turn Rewrite Prompt off and attach an approved output as the reference.
  • Plastic skin or a "rendered" look: remove quality adjectives and add a lens, a light direction and texture words such as "unretouched".
  • Garbled logos, text or brand colours: attach the asset or a swatch instead of describing it, and composite exact marks in post.
  • Characters or products drift across a set: edit from a hero image with a few angle references, and restart from the hero when drift appears.
  • Video subjects morph or the clip drifts from its start frame: one camera move and one action with an endpoint, and describe only motion and change, never the frame itself.
  • Seedance rejects a reference or speaks Chinese: it blocks photoreal human faces in uploads (use Kling O3 for people or a stylised reference), and needs language specified at the end of the prompt.
  • On-screen text is illegible in video: make a still with the text in an image model and animate it as the start frame.
  • A transparent background comes back as a checkerboard: ask for "transparent background" in Auto or choose GPT Image 2; only the GPT Image models produce real transparency.

Still Need Help?​

If you're still stuck, contact support — we're here for you.