GeminiOmni 构建日志

博客

阅读我们最新的产品功能、解决方案和更新内容。

Soundscapes and Skylines: Building Audiovisual Urban Game Scenes with Veo 3.1

How Google Veo 3.1 generates synchronized 1080p video and audio for immersive urban game worlds in a single pass.

2026/08/25
Lena Hoffmann
GPT Image 2 vs Gemini: The Real Cost of AI Video in 2026

GPT Image 2 vs Gemini: The Real Cost of AI Video in 2026

Compare GPT Image 2, Gemini, and FLUX.2 for static assets. Discover why synchronized audio in Veo 3.1 is the missing piece for modern AI video workflows.

2026/08/22
Lena Hoffmann
How to Chat With a PDF Without Losing the Evidence: A Research Workflow

How to Chat With a PDF Without Losing the Evidence: A Research Workflow

A practical PDF chat workflow for mapping a document, asking grounded questions, checking page citations, extracting tables, and writing traceable notes.

2026/08/20
Lena Hoffmann
AI Music Prompt Guide: Describe Structure, Instrumentation, and Change

AI Music Prompt Guide: Describe Structure, Instrumentation, and Change

Learn how to write AI music prompts with clear genre references, instrumentation, tempo, arrangement, dynamics, exclusions, and iteration notes.

2026/08/18
Lena Hoffmann
AI;DR Means Unedited AI Gets Skipped. Here Is a Review Workflow That Survives It

AI;DR Means Unedited AI Gets Skipped. Here Is a Review Workflow That Survives It

AI;DR is not a new summarizer. It is a reader verdict on unreviewed model output. This workflow keeps PDF, image, and video drafts inspectable before they ship.

2026/08/18
Lena Hoffmann
AI Image Aspect Ratios: Choose the Right Canvas Before You Generate

AI Image Aspect Ratios: Choose the Right Canvas Before You Generate

A practical guide to 1:1, 4:5, 3:2, 16:9, and 9:16 AI image aspect ratios, composition choices, cropping risks, and multi-format workflows.

2026/08/13
Lena Hoffmann
AI Image Text Rendering: A Practical Prompt and Editing Workflow

AI Image Text Rendering: A Practical Prompt and Editing Workflow

A reliable workflow for generating readable text in AI images, from short copy and layout prompts to iteration, verification, and post-generation repair.

2026/08/11
Lena Hoffmann
How to Reduce AI Video Flicker, Morphing, and Identity Drift

How to Reduce AI Video Flicker, Morphing, and Identity Drift

A systematic troubleshooting guide for AI video flicker, warped subjects, changing faces, unstable backgrounds, and prompt or source-image drift.

2026/08/06
Lena Hoffmann
AI Video Camera Movements: A Prompt Vocabulary With Practical Examples

AI Video Camera Movements: A Prompt Vocabulary With Practical Examples

A field guide to camera movement prompts for AI video, including pans, pushes, trucks, orbits, crane shots, handheld motion, and combinations to avoid.

2026/08/04
Lena Hoffmann
Image-to-Video Prompt Guide: Describe Motion Without Rewriting the Image

Image-to-Video Prompt Guide: Describe Motion Without Rewriting the Image

Learn how to write image-to-video prompts that preserve the source composition while controlling subject motion, camera movement, timing, and audio.

2026/07/30
Lena Hoffmann
AI Video Storyboard Template: Plan Short Generative Clips Shot by Shot

AI Video Storyboard Template: Plan Short Generative Clips Shot by Shot

A practical AI video storyboard template for turning one idea into coherent shots, prompts, transitions, audio cues, and a repeatable generation plan.

2026/07/28
Lena Hoffmann
使用 Gemini 创建图像的实用指南(2026 版)

使用 Gemini 创建图像的实用指南(2026 版)

Gemini 的图像模型——Imagen 4 和 Nano Banana——能将文本提示转化为成品图像。本文为你揭示文生图的实际工作原理、能奏效的提示词结构,以及最快免费获得首张图像的路径。

2026/06/29
Lena Hoffmann
用 Gemini 制作视频全指南——2026 年简明操作手册

用 Gemini 制作视频全指南——2026 年简明操作手册

Gemini 本身并不渲染视频——这个工作由 Veo 完成,而你需要通过 Gemini 来调用它。本文将详细说明文本转视频的工作原理、真正有效的提示词结构,以及让你快速上手第一条短片的最佳免费路径。

2026/06/11
Lena Hoffmann
我们对 Gemini Omni 的所知——距 Google I/O 大会还有 48 小时

我们对 Gemini Omni 的所知——距 Google I/O 大会还有 48 小时

Gemini 应用中的文本字符串、9to5Google 的泄露信息,以及 Veo 3.1 已具备的功能——它们共同勾勒出 Google 即将发布的产品轮廓。以下是我在大会前一周的解读。

2026/05/12
Lena Hoffmann
我的 Google I/O 2026 主题演讲关注清单——独立开发者真正该听什么

我的 Google I/O 2026 主题演讲关注清单——独立开发者真正该听什么

大多数 I/O 报道会聚焦消费者功能和股价影响。以下是独立 AI 开发者们在 Sundar 于 5 月 19 日登台时应该真正追踪的七个具体信号——它们将改变今年夏天值得构建的内容。

2026/05/11
Lena Hoffmann
Veo 3.1 Fast 为何在独立视频制作中胜过 Sora —— 已公布的按次调用成本

Veo 3.1 Fast 为何在独立视频制作中胜过 Sora —— 已公布的按次调用成本

经过两周的并行生成测试,我确信 Veo 3.1 Fast 是独立视频制作的最佳默认选择。原因并不在于画质,而在于谷歌公布了每秒价格,OpenAI 却将 Sora 隐藏在每月 200 美元的套餐背后。

2026/05/10
Lena Hoffmann
Gemini Live:谷歌2026年最被低估的产品

Gemini Live:谷歌2026年最被低估的产品

实时语音交互,免费套餐中Gemini使用成本为零。所有人都在热切期待Omni;而那个能让你构建皮克斯风格语音助手的模型,自四月起就悄然存在于眼前。

2026/05/09
Lena Hoffmann
Nano Banana 2 与 Imagen 4 —— 何时该选谁(以及我为什么同时支持两者)

Nano Banana 2 与 Imagen 4 —— 何时该选谁(以及我为什么同时支持两者)

两者都来自 Google,都是 2026 年旗舰产品,但各司其职。以下是我使用的决策树,附带有实际提示词和定价,足以证明这一点。

2026/05/08
Lena Hoffmann
百万级 token 上下文窗口并非免费午餐——Gemini PDF 对话的真实页数成本

百万级 token 上下文窗口并非免费午餐——Gemini PDF 对话的真实页数成本

百万级 token 的上下文窗口听起来很神奇,直到你算一笔账。以下是使用 Gemini 2.5 Flash 与一份 1500 页的 PDF 对话的真实成本,以及 GeminiOmni 提供 200 页以内免费聊天的定价逻辑。

2026/05/06
Lena Hoffmann
为什么你的AI工具绝不能在浏览器中存放Gemini密钥——服务端代理完整指南

为什么你的AI工具绝不能在浏览器中存放Gemini密钥——服务端代理完整指南

这个设计模式能保护独立AI开发者免受5000美元意外账单的冲击。本文完整讲解GeminiOmni在每个Gemini API调用前运行的服务端代理机制,附带真实代码模式和边界情况处理。

2026/05/04
Lena Hoffmann
我发布 GeminiOmni 后第一件崩溃的事

我发布 GeminiOmni 后第一件崩溃的事

网站公开 23 分钟后,浏览器里的 Gemini API 密钥就被抓取了。这里有完整的事件经过、损失金额,以及我当天下午立即修改的架构方案。

2026/04/29
Lena Hoffmann