• Deutsch
  • English
  • Alibaba’s Qwen-Image-2.1 combines image generation and editing in a model with publicly available weights that can run on a powerful home computer. At the same time, providers are developing tools for interactive video, automated short-form productions, and cloned voices. Creative Artificial Intelligence (AI) is therefore moving closer to small businesses, creative professionals, and individual users, although licensing, quality, and consent remain unresolved issues.

    One model covers generation and editing

    Qwen-Image-2.1 is an open-weight model, meaning its trained weights, the settings learned during training, can be downloaded. The model has seven billion parameters and combines text-to-image generation, editing of existing images, and work with multiple references. According to the report on Qwen-Image-2.1, it can run on a powerful consumer graphics card such as an Nvidia 3090.

    The practical difference from many earlier tools is that you do not have to switch between a generator and a separate image editor. You can create an image, circle or mask an area, and then request a targeted change. The model can process as many as ten reference images at once, for example to keep people consistent in a group portrait, try on clothing virtually, or design a room using several visual references.

    Another distinctive feature is native support for transparent RGBA images. RGBA is a color format with an additional transparency channel, allowing the model to produce a product without a background or modify text on a transparent layer. A second account of the release confirms that generation, editing, transparency, and support for up to ten reference images are combined in one model.

    The quality claims still require caution. Qwen says its visual component surpasses most closed models on the company’s own benchmark, but independent evaluations are not yet available. Public weights also do not mean unrestricted use: the research license excludes commercial applications, which require a separate license from Qwen.

    Local image AI gives you more control

    A locally operated model processes images on your own computer instead of requiring an external online service. That may be useful if you do not want to upload internal product shots, drafts, or personal photos by default. It does not automatically make setup easy, however: the source identifies a powerful graphics card as a requirement but does not offer a general list of computer configurations that will work reliably.

    For users in Switzerland, this local option is the most directly relevant aspect. It can provide more technical control over where image files are processed, although it does not by itself guarantee complete data protection. The sources offer no information about Swiss support, German-language interfaces, or dedicated availability for Swiss schools and companies.

    “Open” should not be confused with “free for every purpose,” either. You can obtain the model through the listed model platforms and try a demo, but a commercially produced product image or advertisement is not automatically covered by the research license. For freelancers and small businesses, the intended use may therefore matter as much as the model’s image quality.

    Video is becoming a workflow, not just an output

    In video AI, the focus is shifting from individual clips to complete production workflows. ByteDance describes Dramagic as a platform that analyzes scripts, creates characters and storyboards, and turns them into a video preview. Multiple people are supposed to be able to collaborate at the same time, while built-in checks help maintain consistency; access is offered by request through BytePlus.

    One practical example would be a small production team developing a short drama without preparing every scene in separate tools. It could move from script to character designs, storyboard, and preview within one workflow. According to the report about Dramagic, approximately 128,000 short dramas were released in China during the first quarter of 2026, three times the total for the entire previous year; 95 percent were reportedly AI-generated. A Tsinghua University professor estimated the cost of one minute of AI video at $90 to $120, roughly one-tenth of earlier production costs.

    Those figures point to substantial industrial use, but they come from the Chinese industry and expert statements cited in the report. The sector also directly employs 690,000 people. The same report describes allegations that some performers are being forced to transfer their voices and appearances into AI tools before losing their jobs, showing that lower production costs do not automatically produce fair working conditions.

    Runway is pursuing a different approach: videos would begin appearing while you describe them and continue streaming as you refine the request. Instead of entering a prompt, meaning a text instruction, and waiting for a finished clip, you could guide and correct the output immediately. The real-time video generation concept is presented as a research preview, however, not as a generally available finished product.

    Runway argues that faster models could occupy graphics processing units (GPUs) for less time and thereby lower the cost per output. The technical challenge remains substantial because every video frame builds on the previous one: a small error can develop into a major deformation over time. The company says it therefore trains the model on its own outputs as well, helping it correct deviations rather than amplify them, but the source does not provide independently verified results.

    In the same field, OpenAI reports that Higgsfield AI can ship new video advertising features within a day using GPT-6 Astra. According to OpenAI’s account, the technology is intended to make video ad creation easier for small businesses. The brief provider description includes neither prices nor independent quality comparisons, so it indicates a direction more clearly than it proves a practical benefit.

    Voice is becoming a production asset that needs scrutiny

    Voice cloning creates a synthetic voice from a recording of a real person. One comparison of seven application programming interfaces (APIs), technical connections that allow software services to work with other applications, used a ten-second voice sample on each platform. The review assessed similarity to the reference speaker, consent checks, licensing, and cost.

    This makes clear that vocal similarity alone is not enough when selecting a service. For a narrated training video or a recurring audio edition of a company article, the speaker’s permission and the terms for commercial use matter as well. The comparison of seven voice-cloning APIs also considers cost per one million characters, although the available summary does not provide individual prices or identify the winners.

    The connection to image and video AI is straightforward: characters, scenes, and voices can increasingly be created within one production process. This reduces handoffs between tools but also raises the potential for abuse. The risk becomes particularly serious when someone’s appearance or voice is reused without voluntary and informed consent.

    Pros and Cons of creative AI tools

    Pros:

    • Several tasks in one model – Qwen-Image-2.1 generates and edits images, processes references, and supports transparent layers.
    • Local processing – The image model can run on suitable personal hardware, reducing the need to depend on an online service.
    • Faster feedback – Runway’s research approach aims to display video during prompting instead of making you wait after every instruction.
    • Connected production – Dramagic combines script analysis, characters, storyboards, and video previews in one workflow.

    Cons:

    • Unverified performance claims – Independent benchmarks for Qwen’s comparison with closed models are missing, while several other claims come directly from providers.
    • Restricted usage rights – Public model weights do not give Qwen users automatic permission for commercial projects.
    • Errors across video frames – Minor visual deviations can grow as continuous generation progresses.
    • Consent and labor concerns – Cloned voices and digital likenesses can harm people when permission and licensing are not handled properly.

    What this means for your own work

    If you are a beginner, a narrowly defined image project is the most sensible first step. You could remove the background from an existing object, change text on a transparent layer, or use several reference images for a room design. Check the result for visible errors and, before using it commercially, confirm that the license permits your intended project.

    If you already have experience, look at the full workflow rather than isolated outputs. Compare how much time is lost between scripts, storyboards, images, video, and voice, then test tools where they can replace a specific handoff. When evaluating a voice-cloning API, consent checks, licensing, and cost belong alongside speaker similarity.

    More ambitious video projects still require patience. Dramagic is available only by request, Runway’s real-time approach is described as research, and OpenAI’s Higgsfield account provides neither prices nor independent results. “Available to try soon” therefore covers everything from image weights you can already download to features whose general release remains unclear.

    This new category makes creative AI more versatile: images can be generated and edited locally, video processes are becoming more integrated, and voices can enter production through software interfaces. The clearest immediate benefit is less switching between tools and faster experimentation. The main unresolved risks are reliable quality comparisons, commercial rights, and protection for people whose voices or appearances become source material.

    Sources

    AI-FunghiAI-Funghi

    Image: Tranmautritam via Pexels

    © 2024 - 2026 ai-funghi.com | All Rights Reserved | Impressum | Datenschutz