Skip to main content
ALL GUIDES
9 min readUpdated

6 Free AI Audio Tools: How to Actually Use Them

Voice cloning, voiceover, song generation, sound effects and transcription — all MIT or Apache 2.0, all sellable. Try each in a browser first, then one app covers three of them. Includes what ElevenLabs' free tier actually permits.

Verified 18 August 2026. Every link below was opened and checked. Every licence was read from the source.

You don't need to be technical to use any of this. Most of these have a free website where you can try them in about 30 seconds, with nothing to install and no account to create.

Start there. Only install something once you know you want it.


Try everything first, install nothing

These are official free demos. Open the link, use the tool, done. They work on a phone.

What you want to doOpen this
Clone a voiceChatterbox Multilingual
Turn text into a voiceoverKokoro TTS
Make a full song with vocalsACE-Step 1.5
Make sound effectsMOSS-SoundEffect

These run on free shared servers, so you may wait in a short queue. That's the trade for not installing anything.

If you only install one thing, make it Voicebox. It's a normal Mac or Windows app. Download it, double-click, and you get voice cloning and text-to-speech in one window. It has Chatterbox and Kokoro built in, so installing this one app covers three of the six tools below.


1. Chatterbox — clone a voice

What it does. You give it a short recording of someone talking. After that, it can read anything you type in that voice.

Try it free: huggingface.co/spaces/ResembleAI/Chatterbox-Multilingual-TTS — works in 23+ languages.

To install it: don't install it on its own. Install Voicebox (see the top of this guide), which has Chatterbox built in behind a normal app window.

What computer you need. Mac or Windows, and it works with no graphics card at all. There's a smaller "Nano" version built for exactly that, which runs faster than real time on an ordinary laptop.

What usually goes wrong. People upload a noisy recording with music or background chatter, and the cloned voice comes out mushy. Use a clean, quiet recording of one person. Around 10 seconds is plenty.

Can I sell what I make? Yes. MIT licence, no restrictions.

One thing to know. Every file Chatterbox creates carries an inaudible watermark that Resemble AI can detect. You can't hear it and it doesn't limit commercial use, but the audio is detectably AI-generated. If that matters for your work, use a different tool on this list.

Repo: github.com/resemble-ai/chatterbox


2. Kokoro — text to voiceover, no graphics card needed

What it does. You type text, it reads it aloud in a natural voice. It can't copy your voice; you pick from its built-in voices.

This is the answer to "my laptop isn't powerful enough." The whole model is tiny, about 350MB.

Try it free: huggingface.co/spaces/hexgrad/Kokoro-TTS

There's also a version that runs entirely inside your browser, so nothing gets sent to a server at all. Good if you're reading out anything private.

To install it: through Voicebox, where it's one of the built-in voices. There's no official standalone app.

What computer you need. Any laptop, Mac included. No graphics card required.

What usually goes wrong. Only on the manual install route, where it fails because a pronunciation helper called espeak-ng isn't installed. Using Voicebox or the web demo avoids this completely.

Can I sell what I make? Yes. Apache 2.0, no restrictions.

One thing to know. Kokoro hasn't been updated since August 2025. It works fine and the demo is live, but nobody's actively developing it. Don't expect fixes or new voices.

Model page: huggingface.co/hexgrad/Kokoro-82M


3. Applio — turn your voice into someone else's

What it does. You sing or speak into a mic, and it re-voices your recording as a different person. This one changes existing audio rather than reading text, which makes it the tool for covers and character voices.

Try it free: there's no official web demo for this one. You have to install it.

To install it on Windows — genuinely one click, no Python needed:

  1. Download ApplioV3.6.4.zip from the official downloads page. It's about 5GB, so give it time.
  2. Extract it somewhere simple like C:\Applio. Avoid folder names with spaces or accents.
  3. Double-click run-applio.bat. It opens in your browser.

On Mac or Linux, there's no ready-made package and you'll need the terminal. If that's not you, skip this tool.

What computer you need. A Windows PC with an NVIDIA graphics card is the smooth path. It technically runs without one, but slowly enough to be frustrating.

What usually goes wrong. Antivirus quietly quarantines files during extraction, or the folder path has spaces in it. Applio's own docs say to extract to a simple path, temporarily pause antivirus, and don't run it as administrator.

Can I sell what I make? Yes, MIT licence. Small extra step: Applio's terms ask commercial users to email support@applio.org first. Nothing there blocks you, it's just a courtesy notice the others don't have.

One thing to know. The team has announced Applio is moving to security patches only. It's stable and current (version 3.6.4, July 2026), but it isn't growing.

Site: applio.org · Docs: docs.applio.org


4. ACE-Step 1.5 — write a full song

What it does. You describe a song in words, paste in lyrics, and it produces a finished track with sung vocals and instruments. Up to four minutes.

Try it free: huggingface.co/spaces/ACE-Step/Ace-Step-v1.5

To install it — official ready-made packages, no terminal:

  • Windows: download ACE-Step-1.5.7z, extract, double-click start_gradio_ui.bat
  • Mac (M-series): download ACE-Step-1.5.zip, extract, run start_gradio_ui_macos.sh

Python comes bundled, so you install nothing else.

What computer you need. This one has the widest support of anything here: NVIDIA, AMD, Intel Arc, Apple Silicon Macs, and even a plain processor with no graphics card. The small "turbo" model needs under 4GB of graphics memory, which a modest gaming laptop clears easily.

How long a song takes. Under 10 seconds on a decent gaming PC. Much longer with no graphics card, but it does work.

What usually goes wrong. Choosing a model too big for your graphics card, which crashes with an out-of-memory error. Start with the turbo model and work up.

Can I sell what I make? Yes, and this is the strongest position of any music AI right now. MIT licence on both the code and the model itself. The team states it was trained only on licensed, royalty-free and synthetic music, which is a much better place to be than most competitors.

Repo: github.com/ace-step/ACE-Step-1.5


5. MOSS-SoundEffect — make any sound effect

What it does. You describe a sound — rain on a tin roof, a door creaking, a busy café — and it produces that sound at broadcast quality, up to 30 seconds.

Try it free: huggingface.co/spaces/OpenMOSS-Team/MOSS-SoundEffect-v2.0

Honest advice: use the web demo and stop there. This is the one tool on this list with no easy install. No app, no ready-made package, no one-click option. Installing it means the command line, an isolated Python setup, and an NVIDIA graphics card. There's no Mac path at all.

The web demo runs the exact same model, so unless you're generating hundreds of effects, you lose nothing.

Can I sell what I make? Yes. Apache 2.0, no revenue limits, no paperwork.

Model: huggingface.co/OpenMOSS-Team/MOSS-SoundEffect-v2.0


6. Transcription — use Buzz, not faster-whisper

What it does. Listens to a recording and writes out everything said, for subtitles, transcripts or captions.

You'll see "faster-whisper" recommended everywhere. That's a developer library, not an app — there's nothing to click. Install Buzz instead. It's a free desktop app with faster-whisper built in, so you get the same engine in a normal window.

To install it: grab it from the Buzz releases page. Current version is 1.4.4.

  • Mac: Buzz-1.4.4-mac-ARM64.dmg for M-series, -X64.dmg for older Intel Macs
  • Windows: Buzz-1.4.4-windows.exe — Windows will warn you it's from an unknown publisher, which is normal for free software that hasn't paid for a signing certificate
  • Linux: available on Flathub and Snap

What computer you need. Any laptop, Mac included, no graphics card needed. A graphics card makes it much faster but isn't required.

How long it takes. Roughly the length of the recording on a normal laptop. Much faster with a graphics card — 13 minutes of audio takes about a minute on a mid-range gaming PC.

Can I sell what I make? Yes. Both Buzz and faster-whisper are MIT.

A note on the alternatives. MacWhisper is popular but it isn't free past a limited tier, and Pro is €59 to €69. aTrain is free but uses a copyleft licence that's a poor fit for business use. Subtitle Edit is excellent but Windows-only and built for editing subtitles rather than plain transcription. Buzz is the best genuinely-free option that works everywhere.

Repo: github.com/chidiwilliams/buzz


Bonus: Voicebox, the easiest way in

Voicebox is a free Mac and Windows app that puts voice cloning, text-to-speech and dictation in one place. Clone a voice, type text to hear it spoken, or hold a hotkey and speak to type into any app on your computer.

It bundles seven different voice engines including Chatterbox and Kokoro, so it's the single fastest way to get started.

Install: voicebox.sh/download, pick your platform, double-click. About 5 minutes.

What you need. macOS 11+ or Windows 10+, 8GB RAM, 5GB free space. Works with no graphics card, just slower. M-series Macs get a big speed boost from Apple's built-in acceleration. No Linux installer yet.

What usually goes wrong. People think it's broken because the first generation is slow. It's actually downloading the voice model in the background, which can be several gigabytes. Let it finish. On Windows with an NVIDIA card, also go to Settings → GPU and click "Install CUDA backend", or it quietly runs on your processor instead.

Can I sell what I make? Yes, MIT. The app is free and the maintainer has said it stays that way. Paid cloud backup tiers are planned but everything local stays free.

Worth knowing: it's not related to Meta's Voicebox, a 2023 research project whose model was never actually released, despite showing up in "free AI voice" lists constantly. There's also a crypto token attached to this project, which is unusual for an open source app. It doesn't affect using it.


Where to start

Never done this before? Open the four web demos at the top. Nothing to install, works on your phone.

Want it on your computer? Install Voicebox. One download covers voice cloning and text-to-speech.

Need subtitles or transcripts? Install Buzz.

Making music? ACE-Step has a ready-made package for Windows and Mac, and the cleanest licence of any music AI right now.

Need sound effects? Use the MOSS web demo. Installing it isn't worth the trouble unless you're technical.


The part most lists get wrong

All six tools here are MIT or Apache 2.0. You can use every one commercially and sell the output.

That is not true of the tools you'll see recommended most often:

  • ElevenLabs — the free tier is restricted to non-commercial use. Commercial rights start on the paid Starter plan. It's the most-recommended free voice tool and the one you legally cannot monetise for free.
  • Coqui XTTS — the code is open, but the model itself is non-commercial. Worse, the company shut down in January 2024, so there's no commercial licence available to buy.
  • F5-TTS — the code is MIT, the downloadable model is non-commercial.
  • Fish Speech / Fish Audio — released under a research licence. Selling anything made with it requires a separate paid licence, despite being widely listed as free and open.
  • Meta's MusicGen — the code is open, the model is non-commercial. It gets close to 2 million downloads a month and almost none of that output can legally be sold.

The pattern: the code and the model are licensed separately. Sites quote the code licence because it's the one with the friendly badge on the repo page.

Check the model's licence, not the repo badge.

Same homework, done for the other two media: 6 Free AI Image Tools and 6 Free AI Video Tools.