Why most AI tools fall apart on set
The first lesson you learn is that an AI tool’s performance in a lab is meaningless. Whisper transcribes beautifully when you have studio audio. But try running it on set dialogue at 2 a.m. with actors still in frame and wireless mics failing, and you get gibberish. The tool didn’t break. The conditions changed, and the tool was built for ideal conditions only. Production is never ideal conditions. Production is chaos, and chaos reveals every weakness a tool has.
Production designer David Wasco works with strict asset timelines—visual reference needs to be compiled and organized in days, not weeks. He’s tried five different image-organization AI tools. All of them failed under his workflow because they were designed for personal use, not for teams moving at sprint speed. They’d crash on large batches, lose context between sessions, or generate recommendations that looked good but didn’t serve the production’s needs. A tool that works for a freelancer kills you on a 50-person crew because it wasn’t built to think in team time.
Which AI systems actually scale on set
The tools that survive are the ones built with failure modes in mind. When you’re editing for air at midnight, you need a tool that’s either bulletproof or tells you it’s failing before it ruins your export. DaVinci Resolve’s AI features are narrow by design—they don’t try to do everything, they do specific things (color correction, noise reduction, audio sync) and they do those things even when your system is running on fumes. The interface is predictable. The output is predictable. You know what will break and when, which means you can plan around it.
Likewise, Descript works on set because it was built for creators who need transcription, editing, and delivery in the same session. It has built-in redundancy—if cloud transcription stalls, you can still edit locally. It knows that your upload might fail, so it queues work, doesn’t drop it. More importantly, it was designed by people who’ve lived through production stress. The tool doesn’t assume you have perfect audio, perfect wifi, or perfect time. It assumes everything is broken and builds grace into the system. That architectural philosophy is what separates tools that survive deadlines from tools that shouldn’t be trusted near them.

The tools production teams actually use daily
After the liftoff chaos, what separates teams that ship from teams that fracture is the toolset they’ve standardized on. It’s usually unglamorous. Notion or Monday.com for crew coordination, not because they’re AI-powered but because they don’t require your brain to hold ten parallel conversations. Google Drive for asset sharing, because it’s boring and it works. Runway for quick visual effects work when you need it, not instead of a VFX team. RunwayML specifically survives production because it was built to output fast-enough results for temp work and feedback, not final deliverables. It knows its limits and tells you.
The real production AI stack looks like this: a boring, reliable transcription layer (even Otter.ai works if you’re just capturing notes), a predictable editing environment (DaVinci, Premiere, Final Cut), and limited AI assistance in specific places where it’s been tested a thousand times before. Most teams try to apply AI everywhere at once, which is how you end up with a render that fails at frame 47,832 at 4 a.m. Production is not the place to find out where a tool’s limitations are. If you’re going to use an AI tool on set, it needs to have already failed spectacularly at someone else’s production, and someone else needs to have already fixed it.
What every production team needs to test before they deploy
The production teams that avoid disaster run what amounts to a pre-production audit on their tools. They test every AI system under the exact conditions it will face during production: on the actual network speed, with the actual file sizes, using the actual footage or data. If you’re planning to use AI for color grading, you grade a sample reel before you commit. If you’re using transcription, you record test audio from your actual location and see if the tool handles it. If you’re using asset management AI, you load a sample of your actual asset library and push it to breaking point. This takes time you don’t think you have, but it saves the time you won’t have when you’re three weeks in.
Cinematographer Roger Deakins talks about testing every gear choice before production, even if he’s used the gear a hundred times before, because productions vary. The same logic applies to AI tools. What worked last project might crash on this one because your data is different, your network is different, your timeline is different. One production team I know built a 48-hour test sprint before principal photography—they set up the exact tools they planned to use, brought in crew, simulated the real workflow, and discovered that their asset management AI couldn’t handle the image-per-second throughput their cameras would produce. They found this out before day one. They built a hybrid system. Production went smoothly.
Share your thoughts