Blog

  • Some thoughts on AI token arbitrage

    When you are in the tech or AI bubble you surely heard about the big acquisitions announcement of payment infrastructure provider stripe acquiring the AI routing provider openrouter. When you are not part of this bubble the new is new and you probably also won’t ever heard of openrouter. So what makes this company special that stripe put 7 billion dollar on the table to acquire them: openrouter is currently the biggest AI gateway and router provider in the world. There service is routing request to AI models trough their infrastructure and optimize for speed, price or quality. So you don’t need your ChatGPT or Claude subscription any more, you buy credit there and can choose from hundreds of different AI models based on your needs.

    So it is basically when using old economy terms a marketplace for AI tokens. Tokens is the unit in which AI compute is measured and the key underlying economic relevant unit when speaking about AI usage etc. Some say tokens are the new oil or tokens are the currency of the AI boom. So with this all be true, there is an economic effect that is all over the world and fueled our wealth gain through the process of globalization: arbitrage.

    Arbitrage means buying an asset in one market and immediately selling it in another market at a higher price to make a risk-free profit

    AI token arbitrage is something that is not very often spoken about already BUT it exists. In a field where I am active with faktry, there is immense competition that is fought out via extrem price dumping strategies. Competitors offer free of charge access to frontier video models that would cost them normally thousands of compute to run. Of course they are burning VC money and try to buy in user into their subscriptions, BUT there is also arbitrage happening when they buy their compute in cheap cloud regions and sell the access to it in a different country. So here is the classic form of arbitrage happening: buying the the tokens cheaper, than I sell it.

    That’s the basic law of economy and like every profit orientated business should operate exactly like this, you would think, and you are 100% right. But what I fear and what recent deals or new launches I observe make me skeptical about, is that there is a secondary layer of token reselling infrastructure and companies building up that could harm the AI boom and wealth efffect for everyone. When financial institution launch their own AI gateways/routers like Ramp did with router.com this just signals me that old money is using AI tokens as investment leverage that has no foreseeable outcomes for consumer prices. Token or compute could be concentrated, which means prices go up or this secondary layer sits just artificially in between compute provider and consumer taking their arbitrage out of the value stream and increase prices as well for them.

    OpenAI was founded to make AI accessible to basically everyone. This is a very noble vision, but just with creating the models and provide access to it won’t stop opportunistic stakeholder to make use of AI token arbitrage on scale and put the most likely to happen wealth gain for people through AI at risk.

  • Launching faktry – an AI first content and media engine

    This is something big for me, because since months I have been building something really big and challenging in stealth and hardly anyone told about it. Which changes today, as I am publicly announcing my newest venture faktry.

    What Is faktry?

    faktry is a professional-grade, web-based media production platform. It combines over 60 media processing operations with direct access to every frontier over 160 AI image, video, audio, and 3D model – all under a single credit-based account.

    Instead of paying for six tools, juggling five tabs, and managing a separate account for every AI provider, teams and solo creators run everything from one place: image, video, audio, PDF, and 3D, on one shared credit balance. Go from raw assets to polished, brand-consistent output without switching tools or reconciling invoices.

    The core idea: creative work gets scattered across separate generators, editors, and provider accounts. faktry collapses all of it into one workspace.

    Detailed feature overview

    The original idea was to have a simple media processing API i can use with AI automation to process content easy and fast. This was the base and is still a component of faktry. The next step was to create a frontend app for the media processing API and add a full media driven AI gateway that covers now more than 160+ different AI model covering all current needs in terms of content and media creation.

    Most AI tools are a text box wrapped around a single model. faktry ships complete, standalone apps in the same workspace instead – no exports, no round-trips, no extra logins:

    • Creator — A chat-style composer for rapid, high-volume generation: switch between image, video, and audio without leaving the page, chain any result straight into the next generation, and pull in templates or past history with one click. Built for producing many variations fast, not one careful edit at a time.
    • Video Editor — Script-to-scene workbench: write a script, generate per-scene images and video with multiple models, and assemble a storyboard, all in one view.
    • Canvas Editor — Infinite design canvas for compositing images, video stills, text, and shapes, with AI operations (edit, inpaint, upscale, sketch-to-edit) on layers.
    • AI Writer — Long- and short-form content generation with markdown output, inline editing, and brand-voice integration.
    • Node Workflow Builder — Visual node-graph editor with 50+ nodes for multi-step, automated media pipelines.

    Every app shares one workspace, one content library, and one credit balance.

    Platform Capabilities

    faktry AI Gateway

    • One unified interface for 160+ AI models across image, video, audio, document, and 3D – OpenAI, Google, fal.ai, Black Forest Labs, ByteDance, Kling, Alibaba, ElevenLabs, and more.
    • No per-provider accounts, keys, or billing. Requests route to the best model for each task.

    Brand System

    • Multiple brand profiles with logos, color palettes, typography, taglines, and audience definitions.
    • Brand context is injected into AI generation prompts automatically for on-brand output.

    Content Library

    • Centralized asset storage. Search, preview, and re-use outputs across every tool. Saves also your prompt collection

    AI Use Cases & Prompt Gallery

    • Pre-configured operation templates (professional headshots, logo generation, infographics, product photography) with one-click load.
    • Community-contributed and curated prompt library for image, video, and audio.

    MCP Server & API Access

    • MCP server exposes every faktry operation to Claude Desktop, Claude Code, Cursor, Windsurf, Cline, and VS Code – use any operation directly from your IDE or desktop AI assistant.
    • Full API with keys management dashboard, and usage tracking. Designed to work seamless with automation tools like n8n, make or zapier

    How faktry is different from competitor apps?

    DifferentiatorCompetitorsfaktry
    Model breadthSingle provider per toolOpenAI + Google + fal.ai + BFL + ByteDance + ElevenLabs in one place
    Operation depthAI-only or traditional-only65+ AI + FFmpeg/ImageMagick operations in one platform
    Format coverageOne or two formatsImage, video, audio, PDF, and 3D under one roof
    InterfaceA tool, or a single editorFull apps, not forms – video editor, canvas editor, AI writer, and a fast-iteration composer in one workspace
    Brand consistencyNo brand layerNative brand profiles injected into every AI call
    AutomationManual per-file workNode workflow builder for multi-step pipelines
    IDE integrationWeb UI onlyMCP server – use faktry from Cursor, Claude, and more
    BillingA subscription per providerOne credit wallet across every operation

    The way forward

    faktry is a rather big application already with quite some features but there is still room for new stuff: I am in the process of creating a very tight agentic integration with a lots of the features and also having an own faktry creative director agent, that is capable of doing given tasks on it own with the help of the faktry feature set.

    Also I am building a desktop app as a stand alone solution because this is the location where the creative & content work will also in the future happen. AI will not make the file-system obsolet – in fact all the big AI labs have desktop apps with their agents in place with deep file system integration. This is also one of the core concept of the upcoming faktry desktop app.

    The biggest topic for me is to get faktry out of a solo project into a serious business with more people involved, a real tech team and a serious go to market plan to really scale the app. I believe the potential is there and I will try to leverage it.

  • Why I built fleeknote: my own note taking & personal knowledge management app

    Why I built fleeknote: my own note taking & personal knowledge management app

    I am some hooked since years to the topic of knowledge management on a personal and business level. My first university program was a specialization on information and knowledge management in the beginning of the 2000s. Since that a lot changed in the scientific field and also in practice. Web2.0 came and left again, social media shifted into algorithmic media and AI put everything before known to the challenge.

    In 2025 when vibe coding officially got traction I also tried various things and one of the bigger try-outs was a replacement of Google Keep for my personal usage. I started to work with bolt.new on a prototype and worked until mid 2025 every time to time to scale it to a full featured note taking app. At that time also Omnivore (my desired reader app) closed its doors, so I decided to also put the reader features into my own new app.

    I even coined a fancy name for it:

    fleeknote was my own approach to have all in one app for notes (similar concept to Google Keep), with a better task management interface and abstraction, a built in articles reader, a journal and on top built AI features like recalling, summary and an AI agent.

    Below some screenshot of the task management abstraction layer and the article reader:

    Why I built all this?

    Simple answer: because AI and vibe coding enabled me to do this and I had a personal need and passion for the topic. Google Keep is a good product but it has a closed ecosystem with no API or possibilities to connect external apps or services. This was my primary motivation to build fleeknote because I wanted to have an API to run automation etc. based on it. And during the creation process I tried to create add more useful feature to create a truly unique personal knowledge management solution.

    What’s next?

    Is it the one I have always search for? No – unfortunately it is not this universal solution I always wanted but we are getting there. Most of those personal knowledge management solution are very good at storing stuff and make it search-able. But this alone to be just search-able i not enough for me, cause you always need this pro-active intent to do so.

    Real knowledge building is done via connecting things, resurfacing them and putting them into perspective. With the smart recall feature in fleeknote, I tried to build similar patterns into my app, but I have to admit this is not the final evolution step, cause this still needs active intervention and control by the user. The next goal for me is to build a system that does recall and connecting tasks complete autonomous with an presentation layer to the user. Stay tuned!

  • 1 year into vibe coding – my conclusions…

    It is exactly today 1 year ago, that this Tweet below started it all: the term vibe coding was coined by Andrej Karapathy and describes the new process of AI enhanced programming.

    Giving the AI sloppy instruction and expect something working back instead of going down the rabbit hole of thinking, planning and executing with actual hard learned skills. You can think about this development what you want, this year changed how software is crafted completely. The tools and AI model to power them evolved very quickly and AI based software development is one of the most successful AI use case of it all.

    I started my journey way back with the first launch of chatgpt end 2022 trying it out for non-complex software tasks.

    From simple forms to complex apps and games

    Writing software was one of the first use cases I tried to achieve with chatgpt and later claude. Looking back I must admit that I was a long period super disappointed what AI could create in terms of working software. The only thing I really was able to get working were some simple form based calculators I used for my consulting business website. Most of them I created with chatgpt and with php als server language (I use wordpress as CMS so php is available there). That worked out OK but I found it difficult to create something really more advanced.

    The first more complex piece of software I created May/June 2024 when Google podcasts shut down. I absolutely dislike the dark interface of all available podcast player and as podcasts are simple based on open technologies (OPML, RSS) I decided to create my own podcast player. I used again chatpgt here but later switched to claude. The app looks like this and is a real simple podcast player that used an exported opml file as data-source:

    The next big step forward was in November 2024 when I started to use the AI first IDE Cursor. The experience was really a huge step forward as you could just work in your normal app and can get rid of the copy/paste look from AI to code editor and back. The most capable model back then was Claude 3.5 and it was also from this perspective a huge step forward to all know before.

    To test its capabilities I decided to create a browser based shot em up game with it. This worked out pretty OK and I used some time to get a working version out for playing. This game was for me a testing ground also besides software as I created the images, animations and music as well with AI. I documented my learning here.

    Online vibe coding platforms

    Begin of 2025 a couple of new players and a new type of vibe coding emerge: online vibe coding platforms like bolt, replit, lovable or v0 started and gained traction among the vibe coders and software newbies. I started to play around with all of the above mentioned and somehow further used v0 and bolt. With v0 I created more business style apps, as this RSS based news reader app with integrated AI functionality:

    I have to say looking back now at the code and the quality of the apps those platform produced back then, it is a nightmare and none of them can be actually used without massive headache. The typical problems that the AI exposes your AI keys in the frontend sound funny now, but this was the case all the time. As well as massive security flaws with database access, no proper user management etc.

    I experienced the downturn of early vibe coding platform generated code myself: I crafted for quite some time (Jan – June 2025) a Google Keep replacement and over the time extended the functionality to a full blown personal knowledge management solution. But to be really able to publish it and let actual user onto the platform I needed a massive re-factoring which was almost a complete rewrite. Now fleeknote can be used by anyone for free at the moment:

    From all the web based vibe coding platform I must say that I also support the popular opinion that lovable is currently the most capable one. I created several really descent looking apps with almost one shot prompts:

    AI enhanced VSCode forks

    The current setup how I vibe code now is with VS Code and either Claude Code (via terminal), Kilo or Cline. We see here as well the next shift of vibe coding towards real agentic operations where the AI is autonomous over a longer period of time handling given tasks.

    Coming back to the VSCode forks – there are plenty of products around and almost every AI company has its own solution ready to try. The all look quite similar with the “normal” VSCode views enhanced with a chat interface and now with an agent orchestration control panel. Also in terms of the model someone can use there is huge progress and development. A big battle arise between closed source models and open models mostly from Chines AI labs. Claude with its models is on SWE benchmarks still leading but K2 or GLM is very close to it as a fraction of cost. So also the topic of cost efficiency is a valid one in this context now.

    Conclusions

    Some now say that AI agents get worse in coding but as long there is fierce competition among the AI labs and that the race is still on, we will still see loads of improvements and also evolutions of what is vibe coding. What we can clearly see now that most of the improvement come from the smart organization of skills, memory and the orchestration of it not the overpowering of models. So the key to success form me is to apply proven tactics from modern software engineering to vibe coding to organize the output process better and get rid of old common AI flaws.

    My personal conclusion is that we already took a huge step ahead since vibe coding emerged and the first tools were on the market. For really getting the most out of the tools you have to have basic knowledge in software engineering and some experience with shipping software. This is inevitable or you will end up in the news with your app “hacked” because you were to stupid to follow the simplest principles in software dev. Also: please don’t fall for all the AI influencers that sell you vibe coding courses or subscribe to skool community, all the knowledge you need to get started is freely available out there.

  • Creating a pseudo3d racing game is bringing vibe coding to its limits

    Vibe coding is undoubtable an impressive new approach to generate software and products and when you are active on social media a lot, you sooner or later stumble upon some tech influencer trying to tell you its the holy grail while passing you their affiliate links for various such bespoken tools. It is a hype, a lot of people want to own easy and loads of money but we should keep realistically all about this.

    I recently did a lot of vibe coding and for me its really fun to return to tech and programming after having written a single line of code since 10 years. But I also wanted to to advance think then build the 10.000 habit tracker, task manager or whatsoever. I decided to dig deeper into the topic of game development using vibe coding. About my first game I wrote here some time ago. My next game project should be a little more advanced in terms of technics: a pseudo3d racing game in the style of Outrun.

    Pseudo3D what?

    Peudeo3D is a projection technic from the late 80, begin 90s that has been developed back then to create 3D like views for first person view games like racing etc. The core is a logic that projects a 3D world into real world 2D screen coordinates. The math behind this is sort of advanced stuff, but there are good tutorials out there explaining everything:

    Tutorial for building the basic render logic
    Tutorial with working game example

    Vibe coding a Pesudo3D racing game

    The tutorials and sample code mentioned are really crucial for the vibe coding approach here, cause as of End 2025 NO AI model gets the pseudo3d projection correct. If you ask Claude to to a pseudo3d racing game it will come up with something like this. At first sight this looks OK as it can create the projection, we can move the car around, it feels smooth etc. It took about 10 iteration to get to this version – so really fast. BUT: Claude was not able to create the correct car physics not following stupidly the road pattern. It was simply not able to get this right.

    I tried the same example with various currently available flagship models, but none of them (End 2025) was able to get it done. In fact the Claude solution was the best we can get.

    So how to get further? Actually quite simple: You just provide a working example code to the AI and ask to implement the logic from there into your code base. This works better, but also needs manual refinement. The programming language doesn’t matter which is really great. Also I used a python example to get my JavaScript game ahead. Unfortunately the AI never understands the params you need to balance to get a actual usable driving physics out of it, so I ended up tweaking a lot manually. Which has the side effect, you somehow begin to understand the math behind everything.

    My Outrun clone: Neon Drift

    As mentioned before my game should be a Outrun clone with procedural created roads, different surface and background types and a arcade style car physics. I created most of the game graphics using GPT-image-1 and Google’s nano banana. The music was created using Suno and the sound effect I got from the open platform https://pixabay.com. In total I must have invested around 100h to get the version you can play now. It has besides the procedural random track also 11 further pre-defined track you can race. Enjoy!

  • Professional Photos generated with AI compared to 1 year ago

    One of the first real life use cases of generative image generation was when approximately one year ago the flux model with LoRa training made a huge impression in the AI and creative community. It was the first time possible to generate AI images that more or less look like the person trained with the LoRa step. And of course everyone put himself in the fanciest or absurdest places. Besides all the nonsense a really sustainable use case was the generation of professional head-shot that can save you real money as there is no need to visit a professional photographer any more.

    I year later we saw a lot of improvements and new models arise that are capable of doing similar things (eg. put a person into random places/situations) without the need of a extra training step. Google nano banana made a huge impression some weeks ago, but similar output were also already possible in spring when openai launched their gpt-image-1 model.

    Here is a real life example I have created using Google’s nano banana model in AI studio:

    {
    “prompt”: “GQ-style portrait of you smiling contently into the camera. use the outfit from the 2nd image and apply it to the person. Transform this portrait into a high-end editorial photo styled for a GQ magazine cover. Use cinematic studio lighting, glossy highlights, and elegant muted color grading. Keep skin naturally with ultra-detailed textures, enhance contrast, and create a luxury fashion aesthetic. Maintain a clean, minimal background with negative space, centered composition, and polished tones. Output in 8k ultra-realistic quality, crisp, vibrant, and cover-ready.”,
    “negative_prompt”: “blurry, low quality, distorted face, text, logos, watermarks, extra limbs, artifacts, overexposed, poorly lit, cartoonish”,
    “style”: “luxury editorial photography”,
    “resolution”: “8k”,
    “composition”: “centered subject, clean background, cover-ready framing”,
    “lighting”: “cinematic studio lighting with soft shadows and glossy highlights”,
    “output_type”: “photo”
    }

    The result

    The result is pretty impressive considering the inputs I gave the model and how the pose, background and clothing was realized. I cannot be fully happy how my face was drawn by the AI – that is not 100% me.

    Comparing to the flux/LoRa images I created ~1 year ago

    Coming back to the professional head-shots I mentioned at the beginning and were a little bit harder to produce in terms of working with the AI. But except for the image in the middle I can clearly say that my face was well drawn this times and it actually feels like me. There were flaws with flux, like the model puts you always a couple of times into the image (see last picture: the guy in the black sweater is again me) but with some prompt magic, those things were also solvable.

    The way forward…

    My impression is that there is a lot of efforts put into from the big AI players to create advanced image and video models which will bring us even further in possibilities. Character consistence is still a topic that also is almost solved in latest model. In contrary it opens up also the possibilities for easy accessible and hard to recognize image manipulation and miss-use of the tools in all kind of ways.

  • The 3 best AI image models head to head comparison

    Recently there has been a lot of development of new and re-freshed AI model for image generation. Google has officially launched their new image model “nano banana” which took the internet in a hype and pushed the Gemini app to #2 in the app store. Shortly after that seedream 4.0 from Bytedance was released also getting enormous buzz as it is able to produce outputs at the quality of Google’s nano banana at a fraction of the cost. Still also one model worth considering, especially when you are like happy that are also European AI companies still in this race, is flux 1.1 from back forest labs.

    I took those 3 models to a test with the same prompt, to see where each model has its weaknesses and advantages. The order of the images in each row: flux, nano banana, seedream

    In the first image row/prompt flux produced a fine looking image with a nice light atmosphere but was not able to generated the bubble ring correctly. The other 2 model did pass this test but the nano banana output looks sort of pale compared to the seedream image.

    In the 2nd test the image from flux look totally artificial and is like this not usable – although it was the only model who made the lettering in the dress correct and readable. The 2 other images look more realistic but again here the nano banana image rooks really pale compared to the output from seedream. Both were able to generate a good depth effect although the seedream image look more appealing. When you want to create shots of people seedream often comes up with Asian looking faces – probably as a result of its training data (seedream is from bytedance – the company behind tiktok)

    Example three shows again similar result: flux looks to artificial in terms of skin texture but has overall a good image composition. The nano banana image looks really pale this time and rather reminds of a dalle-3 like image. Seedream has again an issue with making Asian looking faces without having prompted for it – other then that the image look quite good.

    Performance Comparison Table

    So it really depends actually on the image/composition you want to create that should actually really influence your model choice. This table is a great summary of when to use which model:

    FeatureSeedreamNano bananaBest For
    Portrait GenerationGood saturation, lighting harmonySuperior sharpness, softer colorsNano banana
    Background Replacement (Text-based)Better lighting matchingInconsistent lightingSeedream
    Background Replacement (Image-based)Slight background alterationsBetter consistencyNano banana
    Multi-Image CompositionDetail-focused, half-body shotsFull-body, overall presentationTie (different strengths)
    Text EditingAccurate text placementText positioning errorsSeedream
    Color AdjustmentToo vivid colorsBetter background matchingNano banana
    Angle TransformationMaintains character consistencySmall, less sharp resultsSeedream
    Photo RestorationEffective stain removal, colorizationMinimal restoration capabilitySeedream
    Object RemovalSlightly blurry resultsSharper facial featuresNano banana
    Style TransferColor-accurate conversionMore artistic brushstrokesTie (different styles)
    High-Resolution EnhancementExceptional detail enhancementMinimal improvementSeedream
  • How to: self-host your apps and AI tools like n8n with coolify

    Running and maintain and own server or similar infrastructure was long a tech only privilege as it requires sysadmin know how. With the rise of PaaS services it got much easier to get a product or application shipped. There evolved a new trend of self hosting applications or services rather then buy access from a SaaS provider. In this brief tutorial I show you how to get a self hosted version of the currently super popular AI & automation tool n8n set up on your own hosted server.

    What is coolify?

    Coolify is a PaaS (platform as a service) platform that can run on your own server environment and lets you manage apps, services, databases and even a set of host servers to run your stuff. It is a free self hosted version of popular platforms like Vercel, netlify or also amazon AWS. With those sort of tools you get an user friendly admin interface to manage your services without the need of logging into a console and doing sysadmin task/commands manually. It is in my opinion a very good alternative for non super tech people to get products shipped.

    What is n8n?

    If you did not heard about n8n you are probably outside of the big AI bubble happening at the moment. It is super popular at the moment and lets you easily build AI agents and other workflow automation. It is open source and besides their cloud offering (free + paid plan) it is also available for self hosting the application. Coolify has a pre-configured service template for n8n so it is just a mather of clicks to get it set up and running without complicated sysadmin work to do. And if you are like me situated in Europa you can host it GDPR compliant in Europe.

    Hosting

    As a hosting provider, as I am situated in Europe, I chose an European one: hetzner.de. The company is very reliable and has a long history of data center management. I have used their servers over the last 20 years in almost all of my agencies I worked for. And they have a very nice feature which makes it really easy for non-tech people to set up coolify: they provide a bundled image with the tool.

    Setup the server and n8n app

    Create a new server in the hetzner console (you have to register up-front for a new account if you don’t have one):

    Then choose from the Apps tab the Coolify App. It would be also possible if you install a docker environment but this makes it much more easier.

    when choosing a server some say n8n runs already with less powerful server with 4GB of RAM. I choose for my instance an Intel variant with 8GB RAM as I also run other apps on the server.

    Wait until the server is ready and then access the coolfiy admin panel. Inside coolify create a new project:

    then add a service and choose the n8n variant with the dedicated database:

    in the settings of the n8n service you can then change the URL:

    having done that you can click the “Deploy” button and see the software get installed.

    To upgrade your n8n instance (which is required quite often as it is in heavy development right now) simply stop the service/container and re-deploy it and it automatically pulls the latest version of the software

    Whats more…

    coolfiy has much more to offer as just to install n8n. I just took it as an example as everyone is going crazy about that tool and wants to build AI agents with it. I wanted to outline a cost effective way to have it also GDPR compliant hosted in Europe.

    There are many other AI related apps you can simple one-click deploy: ollama, flowise, anythingllm, langfuse and many more.

    Of course you can also host your wordpress or drupal website with coolify or you own developed (aka now vibe coded) apps via github repository.

  • How-to: Create AI product ads with consistent characters using Google’s new “nano banana” and Veo3 models

    How-to: Create AI product ads with consistent characters using Google’s new “nano banana” and Veo3 models

    Google’s flagship video generating model Veo3 is around for so time and if you are a frequent tiktok or other video social media platform user you for sure stumbled upon the Yeti, Bigfoot or talking baby videos that have been created with Veo3. They feature quite realistic movements and also lipsynch speech and sound effects. In term of AI video generation it is the benchmark at the moment.

    However it got problems with keeping character consistent over more then 1 prompt/video generation. So when you look very closely the Yeti or Bigfoot has variations over videos from the same account.

    A model named “nano banana” was hyping in the AI see as it made appearance on LmArena and was showing excellent results in terms of realism, consistency and quality. It war rumored that Google is behind this new flagship model and last week we got the confirmation: https://developers.googleblog.com/en/introducing-gemini-2-5-flash-image/ . What sets the model apart from many competitors is the ability to keep character consistency and make super fast context aware edits (for the record: similar workflows would be also possible with flux and flux-kontext models). see example of character consistence images of me with the Past Forward tool:

    What AI tools you need:

    • AI Studio from Google for the flash 2.5 image generation
    • Google Flow for the Veo3 video generation (also possible in the Gemini app)
    • (alternatively the google model you need are also available in fal.ai)

    Character

    random character i created using flux 1.1 ultra. (You can use whatever image model you feel best comfortable with, midjourny has obviously the best result still)

    Product

    For testing i chose my previously with AI generated Jack Daniels gummy bears:

    Step 1: Combining in a product scene using flash image 2.5:

    Generate an realistic image like in a advertising campaign of the person in the image provided sitting in a forest in front of a campfire. in the back we can see his tent – he is obviously on a camping trip. the man is eating gummy bears from the bag shown in the other image. the brand and the visual of the gummy bear bag should be clearly recognizable like in a product ad.

    This is the outcome (first try):

    You can easily adapt the scene more to your needs with simple prompts like remove the whiskey bottle, change the sweater color to green etc.

    Step 2: Creating the product video with Veo3

    We take now the image as input reference for the Veo3 video. With the prompt we bring the image to life and add voice over to the video – like in a real advertising. For more advance use case you can also use Json prompting and my tool I especially created for this: https://veo3json.moweco.com/

    A man eating the jack Daniels gummy bears from his bag sitting in the forest in front of a campfire saying: “enjoy real freedom with the new whiskey flavored gummy bears”

    camera: professional like in an advertising campaign, slowly moving towards the man sitting
    light: natural light, evening mood

    sequence 1:
    man eating from the bag and then saying “enjoy real freedom with the new whiskey flavored gummy bears” and smiling. 0-6s

    sequence 2: big product shot of the gummy bear bag on white background. on the right top side we see then a big yellow background insert “Available now”

    Unfortunately Google flow wouldn’t let my upload and use realistic images of people (cause of country restriction in Europe. So you might want to use a VPN or like me, use the fal.ai Veo3 endpoint for the creation (used to work as well with the Gemini app).

    This is my result (1st attempt – I could have created more versions to optimize and also to get rid of the typo in the end frame or to get the product image exactly in the end frame – but just to demonstrate what a first version looks like):

    The result is far from being a real advertisement someone would use, but I just wanted to briefly show the process in general. Especially the Veo3 output needs more refinements to get a real descent result. Also here the character consistency unfortunately breaks. But with some more tweaks I am sure you can get advertising quality like result with those tools.

  • Why I started to write here again after almost 10 years of a break…

    When I was starting to write texts at this place it was called blogging and the term influencer was not even coined (although I was an early tech focused influencer/blogger with mobilepulse.de). In May 2006 the first postings went live and I continued in a good regular pace till around 2014 to inform about project progress and other minor important things. Somehow then in 2015 I completely stopped to post. Why? I don’t really know, maybe I lost interest in sharing news, got bored of the way of sharing or of the writing process. Honestly I cannot recall 100% why I stopped 10 years ago but I it was a wave of blogs that went silent in that time. Social Media especially social networks took over completely online consumption and writing a blog somehow got out of fashion, I guess.

    First visual appearance of this blog/website

    But there are 2 aspects that made me again write: In 2022/2023 I traveled for a couple of month South East Asia and South America and used newly created pages here to document my travels. It was some sort of travel diary and was great fun to write. From this on I started to write again on a regular basis. The topic I mostly covered was definitely AI and my journey to work with AI. This is very exiting to write as it also documents the progress of AI in the past 2 years. Why I do write this things down? It is a good way for myself to document, reflect and experiment with tools, tasks and approaches. And there is hopefully also value in it for the ones that read the articles.

    What’s next? I don’t know, maybe I will again loose motivation for writing here or this will go on for some time. We will see…