Sound · sepal/audiogen
AudioGen — sound for video, games and podcasts from your description
AudioGen generates sound effects from a text description: footsteps in snow, a creaking door, a thunderstorm, an engine, a button click. In limko the model works in the sound effects generator — you write what should be heard, set a duration from 1 to 10 seconds and get a WAV file. A sound usually takes about 20 seconds and costs 1.5 ₽.
from 1.5 ₽ per sound
What AudioGen can do
- Sound from a description: footsteps, doors, weather, vehicles, animals, interface signals.
- Duration from 1 to 10 seconds — the model does not go longer.
- A WAV file that drops straight onto a track in a video or audio editor.
- Usually about 20 seconds per sound; credits are charged only for a finished result.
- A second opinion right next to it: give the same description to TangoFlux and pick the better result.
AudioGen or TangoFlux
In the sound effects generator the model is chosen before the run, and TangoFlux sits next to AudioGen. That model is cheaper — 0.2 ₽ per sound — and faster: usually a few seconds.
AudioGen is a second take on the same task. These are different models, and they may read the same description in their own ways, so for an important sound it makes sense to run both and compare by ear.
A convenient order: first wording tests on TangoFlux, then the same text on AudioGen if you want another variant.
Sound description examples
Write in your own language: the service translates the sound description into English on its own — sound models were trained on English captions. Short, specific phrases work best: source, material, setting.
Footsteps
Footsteps in heavy boots on fresh snow, an unhurried rhythm, a crunch with every step, a quiet winter street.
Door
An old wooden door slowly creaks open and slams shut with a dull thud, an empty hallway with a slight echo.
Weather
Heavy rain drumming on a tin roof, thunder rolling in the distance.
Interface
A short, soft button click in a mobile app, clean sound with no echo.
Where sound effects come in handy
A generated effect saves the day when the sound you need is not in a library or has to sound in a very specific way.
| Where | Which sounds |
|---|---|
| Video and editing | footsteps, a door slam, street noise, a passing car |
| Games | the hero’s footsteps, a creaking chest, a beast’s growl, rain on a level |
| Podcasts | rain in the background, a phone ringing, a transition sting between topics |
| Presentations and interfaces | a click, a notification, an error sound |
Frequently asked about AudioGen
How is AudioGen different from TangoFlux?
TangoFlux is cheaper and faster, while AudioGen is a different model for the same task. They may respond differently to the same description, so an important sound is worth trying on both.
Can I make a sound longer than 10 seconds?
No, AudioGen makes sounds from 1 to 10 seconds. A longer ambience can be assembled from several fragments in an editing program.
Can I write the description in my own language?
Yes, write in your own language: the service will translate the description into English, the language the model was trained on. If the sound comes out wrong, be more specific about the source and setting.
How much does one sound cost?
One generation costs 1.5 ₽. You pay with credits, no subscription, and if a generation fails, the credits return to your balance automatically.
What format is the file?
WAV — you can drop it straight onto a track in any video or audio editor.