Skip to main content
ai-audio

AI Sound Effects: How to Generate Custom SFX From Text

Oakgen Team9 min read
AI Sound Effects: How to Generate Custom SFX From Text

An AI sound effects generator turns a written description ("heavy wooden door creaks open slowly in a stone hallway, 3 seconds") into an audio file in a few seconds. Use it for foley, ambience, UI clicks, whooshes and impacts, and look elsewhere for speech or for music that has to hit an exact tempo.

Good results depend on writing the prompt the way a sound designer would, naming the source, the action, the material, the space, the length and the mic perspective.

Research note

On September 26, 2026 we checked model limits and settings using ElevenLabs' API reference and docs, fal's CassetteAI page, and the configuration Oakgen runs in production. Oakgen is our own product, and nothing here comes from a listening test or blind comparison.

Quick answer: which AI sound effects generator to use

JobBest pickWhy
Transitions, whooshes, risers for YouTubeElevenLabs sound effectsExact duration from 0.5 to 30 seconds, so the sound matches the cut
Ambience beds (rain, room tone, city)ElevenLabs with looping, in its own app or APINative loop option avoids an audible seam
Cheap drafts and throwaway one-shotsCassetteAIFlat price per clip regardless of length
Game UI clicks and pickupsElevenLabs at 0.5 to 1 secondShort durations and higher prompt influence keep sounds tight
Dialogue or a voice lineA text-to-speech model, not SFXSFX models do not produce reliable, intelligible words
Anything on a beat (stingers, jingles, intros)A music generatorSFX models do not hold tempo or bar lengths

How text-to-sound effects work

A text-to-SFX model is trained on a large library of labeled audio. When you type a prompt, it generates a new waveform that matches the description. It does not pull a file from a stock library, so every generation is new, and two runs of the same prompt sound different.

What you control:

  1. The prompt. Most of the quality comes from here. Specific nouns and materials beat adjectives.
  2. Duration. Set it and the model fills exactly that length. ElevenLabs will guess a length if you leave it empty.
  3. Prompt influence. A 0 to 1 dial that trades creativity for literal accuracy. ElevenLabs defaults to 0.3 in its API. Higher values follow your words more closely, with less variety.

ElevenLabs' own playground generates four variations per prompt, and it's a good habit on any platform to make several and keep the best.

Where AI sound effects are strong, and where they fail

Text-to-SFX handles these well:

  • Foley: footsteps, cloth rustle, glass set on a table, keys, zippers, paper.
  • Ambience: rain on a window, forest at dawn, office hum, crowd walla, wind.
  • UI and game sounds: clicks, confirms, errors, coin pickups, menu swipes.
  • Transitions: whooshes, risers, swells, reverse cymbals, glitch hits.
  • Impacts: punches, body falls, cinematic booms, braams.

It struggles with:

  • Dialogue. You may get mouth sounds or gibberish that resembles speech. Use a text-to-speech voice model instead.
  • Precise musical timing. ElevenLabs does accept prompts like "90s hip-hop drum loop, 90 BPM", but a stinger that must land on bar four of your track is a music job for a music generator.
  • Complex sequences in one clip. ElevenLabs' guide recommends generating individual effects and combining them in an editor. That advice holds for every model.
  • Very specific real-world sounds. "A 1967 Mustang starting" gets you a plausible old V8 that won't match that specific car.

The prompt formula: source + action + material + space + duration + perspective

Most weak SFX prompts are one or two words ("explosion", "door"). Write six parts instead:

PartQuestion it answersExample
SourceWhat makes the sound?Heavy oak door
ActionWhat is it doing?Creaks open slowly, then latches
MaterialWhat is it made of, or hitting?Iron hinges, stone floor
SpaceWhere is it?Long stone hallway with echo
DurationHow long?3 seconds (set in the duration field too)
PerspectiveWhere is the mic?Close-up, dry / distant, reverberant

Put together: "Heavy oak door with iron hinges creaks open slowly and latches, long stone hallway with echo, close-up, 3 seconds."

A few rules that save generations:

  • Use sound design vocabulary. ElevenLabs' docs list the terms its model understands, including impact, whoosh, ambience, one-shot, loop, stem, braam, glitch and drone.
  • Say "one-shot" for single hits and "loop" or "ambience" for beds.
  • Describe what you hear, not what you see. "Villain enters" means nothing. "Low sub drone with metallic scrape" does.
  • Match duration to the edit. A 10-second whoosh for a 0.5-second cut forces you to trim and fade.

45 AI sound effect prompts, grouped by job

Copy these as a starting point, then change the material and space to fit your scene. Durations are suggestions; set the same number in the duration field.

YouTube transitions (8)

  1. Fast airy whoosh passing left to right, clean, no reverb, 1 second
  2. Deep cinematic whoosh with low sub tail, 2 seconds
  3. Reverse cymbal swell rising into a soft hit, 2 seconds
  4. Short digital glitch hit with bit-crushed texture, 0.5 seconds
  5. Paper page flip, crisp and close, 0.5 seconds
  6. Camera shutter click with film advance, 1 second
  7. Tape stop slowing down to silence, 1.5 seconds
  8. Rising tension riser with white noise and pitch bend up, 4 seconds

TikTok and Shorts (6)

  1. Cartoon boing spring bounce, bright and playful, 1 second
  2. Record scratch, vinyl, sharp and dry, 1 second
  3. Cash register cha-ching with coin rattle, 1.5 seconds
  4. Comedic slide whistle going down, 1.5 seconds
  5. Crowd gasp, medium room, 2 seconds
  6. Short pop with bubble texture, one-shot, 0.5 seconds

Game UI (7)

  1. Soft UI click, plastic button, close, dry, 0.5 seconds
  2. Positive confirm chime, two ascending tones, glassy, 1 second
  3. Error buzz, low and short, muted, 0.5 seconds
  4. Coin pickup, bright metallic ding with sparkle tail, 1 second
  5. Inventory open, leather bag rustle and buckle, 1 second
  6. Level up fanfare burst, magical shimmer, 2 seconds
  7. Menu swipe, soft air whoosh, 0.5 seconds

Horror (6)

  1. Low sub drone with distant metallic scraping, dark, 10 seconds
  2. Old wooden floorboard creak under slow weight, close, 2 seconds
  3. Faint children's music box winding down, slightly detuned, empty room, 6 seconds
  4. Sudden orchestral stab hit with reverb tail, jump scare, 2 seconds
  5. Heavy breathing behind a closed door, muffled, 4 seconds
  6. Chains dragging across a concrete basement floor, reverberant, 4 seconds

Nature ambience (6)

  1. Steady rain on a window with occasional distant thunder, interior, 20 seconds
  2. Forest at dawn, layered birdsong, light breeze in leaves, 20 seconds
  3. Ocean waves breaking on a pebble beach, medium distance, 20 seconds
  4. Crackling campfire, close, with soft wood pops, 15 seconds
  5. Wind howling over an open mountain ridge, 15 seconds
  6. Summer night crickets and frogs by a pond, 20 seconds

Product ad foley (7)

  1. Soda can opening with fizz, close-up, crisp, 1.5 seconds
  2. Pouring sparkling water into a glass with ice, close, 3 seconds
  3. Sizzling steak placed on a hot cast iron pan, close, 3 seconds
  4. Premium box lid sliding off slowly, soft friction, 2 seconds
  5. Lipstick cap clicking shut, satisfying, close, 0.5 seconds
  6. Sneaker squeak on a polished gym floor, 1 second
  7. Mechanical keyboard typing, tactile switches, close, 4 seconds

Podcast stingers and bumpers (5)

  1. Soft vinyl crackle intro with warm room tone, 3 seconds
  2. Radio dial tuning through static, landing on silence, 2 seconds
  3. Newsroom teletype ticking, 3 seconds
  4. Short bright chime for a segment break, 1 second
  5. Microphone switch click and brief feedback hum, 1 second

For podcast stingers that must feel like part of your theme music, a sound effect is usually the wrong tool. Our guide to podcast intro music covers building a matching bed and stinger from one prompt.

Generate These Sound Effects in Agent Chat

Paste any prompt from the pack, name the length you need, and ask for three variations to compare.

Generate a Sound Effect

ElevenLabs sound effects settings, and how Oakgen maps them

The setting names differ slightly between ElevenLabs' own product and hosts that run its model. Here is what each exposes, checked against ElevenLabs' API reference, fal's CassetteAI page and Oakgen's model configuration.

SettingElevenLabs native (app and API)ElevenLabs SFX on OakgenCassetteAI on Oakgen
Modeleleven_text_to_sound_v2eleven_text_to_sound_v2cassetteai/sound-effects-generator (via fal)
Prompt lengthUp to 450 characters in the playgroundUp to 450 charactersUp to 1,000 characters
Duration0.5 to 30 seconds, or empty to auto-pick0.5 to 30 seconds in 0.5-second steps, default 101 to 30 seconds, whole seconds, default 10
Prompt influence0 to 1, default 0.30 to 1, default 0.5Not exposed
LoopingYes, boolean optionNot available (sent as off)Not available
OutputMP3; WAV 48kHz for non-looping in the appMP3, 44.1kHz, 128kbpsWAV

What to change first:

  • Prompt influence. Push it up (0.6 to 0.8) for UI sounds and foley, where you want exactly what you typed. Pull it down (0.2 to 0.4) for ambience and atmospheric drones, where variation sounds more natural.
  • Duration. Always set it. The auto-length guess is fine for exploring, but a fixed length is what makes a sound fit an edit.
  • Model choice. Use ElevenLabs for anything that ends up in a finished video. Use CassetteAI for rapid drafts, especially long ambience, because its price does not rise with length.

ElevenLabs sound effects: max duration, looping and credits

  • Max duration: 30 seconds per generation. Longer beds need a loop or several clips stitched together.
  • Looping: available on eleven_text_to_sound_v2 in ElevenLabs' own app and API. It is meant for ambience that repeats with no click at the seam.
  • Credits on ElevenLabs: its docs state 40 credits per second when you specify a duration. Those are ElevenLabs credits, which differ from Oakgen credits. For its plan prices, see our ElevenLabs pricing breakdown.
  • Commercial use: covered in the licensing section below.

What AI sound effects cost on Oakgen

Oakgen charges the provider's cost converted at 260 credits per US dollar, rounded up, with no platform markup. ElevenLabs sound effects are priced per second of audio. CassetteAI is a flat fee per clip. These figures are as of September 2026.

ClipElevenLabs SFXCassetteAI
1 second2 credits3 credits
3 seconds4 credits3 credits
5 seconds7 credits3 credits
10 seconds13 credits3 credits
20 seconds26 credits3 credits
30 seconds39 credits3 credits

In plan terms, the 2,000 credits that come with Basic each month pay for roughly 150 ten-second ElevenLabs effects, before you spend anything on images, video or voice. A full 45-prompt pack at the suggested durations, one take each, lands well under 400 credits. Check the pricing page for current plans.

Current Oakgen limitation

Oakgen's audio tool page currently has tabs for voice generation and voice cloning, but no dedicated sound effects tab. Sound effects run through Agent Chat or the Oakgen MCP server, where you ask for the effect and duration in plain language. Looping is also not available on Oakgen yet. If you need native loops today, ElevenLabs' own app is the better tool for that job.

How to make sound effects for YouTube videos: layering under video

Placement matters as much as the sound itself. This workflow holds up in most edits:

  1. Spot the edit first. Watch the cut and list every moment that needs sound: transitions, on-screen actions, text pops, the ending.
  2. Generate one sound per moment. Match the duration to the gap on the timeline. Generate three or four takes.
  3. Lead the picture slightly. Nudge whooshes and risers earlier so they peak exactly on the cut. For impacts, line up the transient with the frame where contact happens.
  4. Build complex moments from layers. A car door slam is a latch click, a body thud and a short room reverb. Three simple prompts usually beat one complicated one.
  5. Duck the music. Drop your background track a few decibels under key effects so they read. Treat voiceover the same way and keep effects below the voice.
  6. Fade everything. Short fades on both ends remove clicks. For a fake loop, crossfade the end of an ambience clip into its start.
  7. Keep a bed running. A quiet room tone or ambience under the whole video hides the silence between edits and makes a faceless video feel less sterile.

If you produce faceless channels, pair this with our faceless YouTube videos guide. For cutaways, the AI B-roll workflow shows where effects make generated footage feel shot rather than rendered. Some newer video models generate their own audio; our roundup of AI video models with native audio covers when that is enough and when you still need separate SFX.

Add a Music Bed Under Your Effects

Generate an original instrumental track at the length of your edit, then layer your sound effects on top.

Generate a Music Bed

AI sound effects licensing: what you can actually publish

The provider and your plan set the licensing terms.

  • ElevenLabs. Its help center says free plan content cannot be used commercially and requires attribution to elevenlabs.io. Paid plans include a commercial license that stays valid indefinitely, excluding beta services and subject to its service-specific terms. Our ElevenLabs commercial use explainer goes deeper.
  • Oakgen. No ownership claim on your inputs or eligible outputs, per Oakgen's terms. To use an effect commercially you still have to comply with those terms, the law, and the policies of the model provider behind it, and Oakgen gives no guarantee that an output is original or non-infringing.
  • Recognizable sounds. Do not prompt for trademarked sonic signatures, such as a famous startup chime or a network jingle. A brand sound belongs to its owner even when a model recreates it. If you want your own, start with our sonic logo guide.
  • Keep records. Save the prompt, model, date and plan for any effect that goes into a paid ad or a client deliverable.

Using effects in ads, jingles and games

The radio ad script template marks where SFX cues go in 15, 30 and 60-second scripts. A brand running radio or streaming ads usually pairs a voice read, a music bed and a sonic logo with a handful of signature effects that repeat across spots. If the ad needs a sung hook, see how to make a jingle. Game developers building a full soundtrack can follow our indie game background music guide and use the UI prompts above for the menus.

Voice the Script That Goes Over Your SFX

Write the narration, pick a voice, and generate it next to your effects so the whole mix comes from one credit balance.

Generate a Voiceover

Sources

AI sound effectssound effects generatorElevenLabs sound effectstext to sound effectsSFX promptsYouTube sound effectsfoleyaudio for video
Share

Related Articles