HN
Today

Show HN: Vocal Slice – Cut audio by selecting text, fully on-device

Vocal Slice is a new desktop application that radically simplifies audio editing by allowing users to cut clips simply by highlighting words in an AI-generated transcript. This 'Show HN' highlights a clever application of on-device AI (Whisper) to solve a common, tedious problem for podcasters, video editors, and content creators. Its local processing ensures privacy, a feature highly valued by the Hacker News community for sensitive projects.

33
Score
20
Comments
#10
Highest Rank
6h
on Front Page
First Seen
Aug 17, 5:00 AM
Last Seen
Aug 17, 10:00 AM
Rank Over Time
171510101415

The Lowdown

Vocal Slice is a novel tool designed to streamline the audio editing workflow, specifically for tasks involving voice recordings. It allows users to intuitively cut audio segments by interacting directly with a text transcript, rather than meticulously scrubbing waveforms. The application performs all its core functions, including transcription and slicing, entirely on the user's device, ensuring privacy and offline capability.

  • AI-Powered Transcription: Utilizes local Whisper models to transcribe audio files (WAV, MP3, FLAC, M4A, AAC, OGG) with word-level timestamps.
  • Text-Based Editing: Users highlight phrases in the transcript, and the waveform automatically adjusts to that precise segment, with fine-tuning handles for accuracy.
  • Lossless Export & Naming: Exports byte-perfect WAV slices from original sources, preserving quality, and allows for templated file naming conventions.
  • Re-trimming Capability: Existing slices can be re-opened and their boundaries adjusted without needing to re-process the audio.
  • Multilingual Support: Supports various languages through different Whisper models.
  • Performance & Privacy: Benefits from GPU acceleration (WebGPU) and guarantees no audio data ever leaves the user's machine, making it suitable for sensitive or NDA-protected content.

This innovative approach promises to significantly reduce the time and effort traditionally involved in pulling clips from long audio recordings, offering a modern, efficient solution for anyone working with spoken audio.

The Gossip

Transcription Truths and Tool Talk

The accuracy of Vocal Slice's local transcription was a key point of inquiry. Users questioned if the text-to-speech component was '100% accurate,' a critical aspect for a text-based editing tool. The developer clarified that the tool uses local speech-to-text to record word-level timestamps, which are then used to precisely align text selections with audio in-and-out points, explaining the mechanics behind the seemingly magical text-to-audio cutting.

Pricing Predicaments & Proprietary Probes

A significant portion of the discussion revolved around the monetization strategy and the project's open-source status. Commenters questioned the annual subscription model versus a one-time purchase. The developer defended the subscription, citing plans for future feature additions like 'um and ah' removal and advanced editing capabilities. The tool was confirmed as proprietary, though transparent about its local, AI-driven functionality. There were also technical queries about using GitHub Releases for proprietary binaries and a missing license file, which the developer committed to addressing.

Versatile Visions and Vocal Ventures

Users enthusiastically explored diverse potential applications and future enhancements for Vocal Slice. Suggestions ranged from practical uses like preparing audio for YouTube videos and creating 'YouTube poops,' to more sophisticated needs like assembling long-form audio drama from multiple takes. The developer confirmed the tool's utility for creative cuts and even highlighted its existing (though under-advertised) ability to quickly navigate and select between multiple instances of the same spoken line, and expressed plans for native video support.