Ask HN: Are there AI models for generating sounds based on a text and reference?
An 'Ask HN' post probes the frontier of AI audio generation, seeking a model that can synthesize sound effects from both text prompts and a reference audio input. The author highlights the stark contrast between the maturity of image and text-to-audio models versus the apparent immaturity of multimodal audio generation, sparking curiosity about this niche's current state.
The Lowdown
A Hacker News user is searching for an artificial intelligence model capable of generating new audio, specifically sound effects, by combining a text description with a reference audio input. Despite the prevalence of sophisticated multimodal AI for image and text outputs, this specific Audio + Text -> Audio workflow appears to be a challenging or underserved area.
- The user notes that
Image + Text -> Image(with reference image guidance) andAudio + Text -> Text(for descriptions or classification) are relatively mature and commoditized problems in AI. - Their use case involves generating new sound effects, particularly when clean examples for a source are scarce, hoping to leverage existing audio alongside text instructions.
- They have attempted solutions like ElevenLabs SFX, which only offers text-to-audio generation without reference input, making accurate sound recreation difficult.
- Stable Audio 3, advertised as providing the desired functionality, yielded 'terrible' results in their experience.
- The author concludes that this particular AI modality seems to be significantly behind other areas, comparing its current state to image generation technology from 2016.
The core question posed is whether such a multimodal audio generation solution exists commercially, or if it represents a less common or currently unsolved problem in the AI landscape.