Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Making Things With AI

Images, video, voice and music — how they work, where they break, who owns them.

Lesson 62 of 848 min

Sound effects, foley and where it is already good

The one place generation is unambiguously ready

Everything in this module so far has been qualified. Sound effects are the exception.

A sound effect is typically under five seconds, has no long-range structure, no polyphony to keep independent, and no requirement to be a specific recognisable thing. It is texture and envelope. That is precisely the regime where generative audio is strong.

The practical consequence: a text-to-sound model produces usable effects for a large fraction of what a video needs — impacts, whooshes, ambiences, mechanical noises, footsteps on surfaces, weather, crowd tone, interface sounds. In seconds, at any length, tuned by re-rolling.

Why this matters more than it sounds

Sound is half of video, and it is the half amateurs neglect. A well-shot clip with no sound design reads as amateur; an average clip with good sound reads as professional. This is not an opinion, it is a reliable finding of anyone who has run the comparison.

Historically, the barrier was libraries. A decent effects library cost money, and the free ones were small, badly labelled and inconsistent. Generation removes that barrier for a large category of sounds.

There is also a real advantage over libraries, and it is worth naming: an effect can be generated to length. A three-second whoosh that must be 2.4 seconds is a fight with a library file and a re-roll with a generator.

Where it still fails

Specific real objects. A particular model of car, a named machine, a specific bird. The model produces something in the category, and anyone who knows the object will hear that it is wrong.

Speech-adjacent sounds. Crowds where words are almost intelligible, laughter, a baby crying. These sit close enough to speech to trigger the model's speech machinery and come out uncanny.

Musical instruments in isolation. Better handled by a sample library or a real instrument.

Anything that must sync tightly to picture. Generation produces a sound, not a sound with a hit point at 1.32 seconds. You place it in the edit, which is normal practice anyway.

The free path, which is strong here

This is a category where free sources are genuinely excellent and deserve to be tried first:

  • Freesound — a large community library, mostly Creative Commons, with licences stated per file. Check each one; they differ, and some require attribution.
  • The BBC sound effects archive, available for personal, educational and research use under its own terms.
  • Several national archives and museum collections with open licences.
  • Recording it yourself. A phone in a quiet room records better foley than most people expect. Footsteps, cloth, doors, water, paper — all of it is available in your own house, it is free of any rights question, and it is more specific than anything generated.

That last one deserves emphasis because it is chronically undervalued. Foley artists have always made sounds with objects rather than recording the real thing, because the real thing usually sounds wrong. This craft is available to anyone with a phone and a quiet hour, and it produces the most convincing results of any option here.

The workflow

The professional shape is a layered one, and it applies whatever the sources:

  1. Ambience — a continuous bed establishing the space. Generated ambiences are excellent for this.
  2. Spot effects — the specific events. Generated, library, or recorded.
  3. Detail — the small sounds that sell it: cloth, breath, a distant car. This layer is what separates good sound from adequate sound and it is almost always missing from amateur work.

Mix these in any free editor. The rule that matters more than the sources: quiet detail under a scene does more than loud effects on it.

One technique is worth knowing because it costs nothing and improves almost any sound: layering. A single generated impact rarely sounds convincing on its own. The same impact plus a low thump underneath and a short bright crack on top does, because that is how real sounds are built — a body, a low end and a transient. Generate three variations, stack them, adjust the relative levels and trim the starts to line up. This one habit converts adequate generated effects into good ones and is the reason a sound designer's library of individually unremarkable files produces remarkable results.

The honest limitation: generated effects are unspecific by nature. For a documentary about a particular factory, generated machine noise is a fabrication, and using it raises the same disclosure question as any other synthetic element in factual work. For drama, atmosphere and explainer video, it is simply a tool, and a good one.

The one thing to keep

Generated sound effects are at production quality for many categories because a sound effect is short and unstructured, which is exactly the regime these models handle well.

Before you move on

Why are generated sound effects at production quality when generated music is not?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly

Sound effects, foley and where it is already good · Making Things With AI · Addaly