Controlling "Zoom" With Words Alone
How to make an image model frame your subject the way you want, from tight portrait to roomscale, using nothing but prompt structure.
The guide is mainly for Krea2, but it will probably work just as well for other models.
Sidenote: Article was written by me, and corrected by Fable for grammar and formating.
The problem: everything becomes a close-up
Ask an image model for "an e-girl in her gaming room" and you will usually get a centered, chest-up shot with a blurry hint of a room behind her. That is not a bug, it is what the training captions taught the model. Most photo captions describe the person first and the place as an afterthought, so the model learned: whatever the text spends its words on is what fills the frame.
That is also the key to controlling it. "Zoom" is not a camera setting you can name, it is a budget. The frame belongs to whatever you describe first, longest, and in the most detail. Once you treat word order and word count as your zoom ring, framing becomes controllable and repeatable.
The technique: order and word budget are your zoom
One lever runs the whole thing: what you describe first, and how many words you spend on it. Whatever gets the words gets the frame. That gives you two directions.
Person first → zoom in. Open on the face and body, spend your words there, and let the room trail off as a blurred afterthought. The model fills the frame with the person.
Room first → zoom out. Open on the space with walls, floor, furniture, light etc, and only then drop the person into it. By the time she shows up, the room already owns most of the frame, so she shrinks to fit what is left.
Everything between a tight portrait and a wide roomscale is just how far you slide between those two: how early the person appears, and how big a share of the words she gets.
Tip: if you write your prompts with an LLM: you don't have to build this ordering by hand. Add an instruction like "Start by describing the scene around the subject, establishing the room in detail first. Then place the persons inside the room, seen at a distance." and the model applies the room-first structure for you. Adjust that line to slide the framing wherever you want.
Examples: the framing ladder
Same girl, same room, four framings.
A 22-year-old Korean woman with pink hair, an oversized hoodie, a black RGB-lit room. The only thing I move is the words.
Each prompt is two parts: a subject part and a scene part. Watch two things happen as you go down:
The balance of words tips from the subject toward the scene.
At the wide end, the scene moves ahead of the subject and is described first.
That combined shift is the entire zoom control. Nothing else does the work.
1. Portrait — the subject is the whole prompt
A close photo portrait of an 22-year-old korean woman, face to the camera: playful half-smile, winged eyeliner, glossy lips, soft pastel-pink hair, oversized gaming headphones glowing pink on the earcups.
Behind her, just a blur of pink LED light.
Key moves:
Nearly every word describes the face and head
The room is a single blurred clause, no detail
Heavy face detail pulls the camera in tight

2. Two-thirds body — subject leads, the room appears
A three-quarter photo shot of an 22-year-old korean woman, framed from just above her head down to her thighs, lounging sideways in a gaming chair. Face to the camera, pastel-pink hair, an oversized hoodie slipping off one shoulder, pink headphones.
Behind her, a black bedroom in soft focus: RGB strip lights and the glow of a monitor.
Key moves:
Explicit crop: "from just above her head down to her thighs"
A pose (lounging in the chair) anchors the body
The room enters, but only as one soft-focus line

3. Fullbody — the room comes first, then the full figure
A photo of a black gaming bedroom lit by RGB strips: a black-and-pink gaming chair, a desk with a glowing monitor, LED lights tracing the walls, plushies on the bed.
In the chair, full body shot of an 22-year-old korean woman with pastel-pink hair, oversized hoodie and thigh-high socks, smiling cute, her feet touching the carpet.
Key moves:
The room is described first, before the person
"full body shot" states the framing outright
The face is kept to two words ("smiling cute")
Describe feet to force the feet to be in frame

4. Roomscale — the room leads, the person sits in the distance
A wide view photo of a whole black gaming room: RGB-washed walls, a desk with dual glowing monitors, shelves of plushies, string lights, a neon sign, and a bed against the back wall.
Sitting on the edge of the bed, an 22-year-old korean woman with pink hair, smiling to the camera in the distance.
Key moves:
The whole first paragraph is the room
The person is one short clause, placed "in the distance"
A placement anchor ("on the edge of the bed") sets the scale since the bed is defined at the back wall

Cheat sheet
Zoom = word budget. The frame belongs to what you describe first and most.
To zoom out: room first, person second, and starve the face of words.
Want a body part in frame? Describe it — don't lean on crop labels. "Standing barefoot on the white rug" pulls the feet into shot even when the framing is tight; naming a crop alone often fails.
Anchor the body to the room — "feet on the floor", "one hand on the desk", "sitting back in the chair". Contact points force those parts into frame and set the scale.
All example images in this article were generated from the exact prompts shown, unedited.
仅用文字控制"变焦"
如何只靠提示词结构,让图像模型按你的意图取景——从紧凑的面部特写到整个房间的远景。 本指南主要针对 Krea2,但很可能同样适用于其他模型。 附注:本文由我本人撰写,由 Fable 协助修正语法与排版。
问题:一切都变成了特写
向图像模型要"一个在游戏房里的电竞女孩",你通常会得到一张居中的、胸部以上的照片,背景里只有一片模糊的房间轮廓。这不是 bug,而是训练图注(caption)教给模型的东西。大多数照片的图注都先描述人物,把场景当作事后补充,于是模型学到了:文字花在谁身上,画面就属于谁。
而这也正是控制它的关键。"变焦"不是一个你能直接点名的相机参数,它是一种预算。画面属于你最先描述、描述最长、描述最细的东西。一旦你把词序和词量当作变焦环来用,取景就变得可控且可复现。
技巧:词序和词汇预算就是你的变焦环
整个方法只有一根杠杆:你先描述什么,以及在它身上花多少词。谁拿到词,谁就拿到画面。 由此产生两个方向。
人物在前 → 拉近(zoom in)。 开头就写脸和身体,把词都花在那里,让房间只作为模糊的收尾一带而过。模型会让人物填满画面。
房间在前 → 拉远(zoom out)。 开头先写空间——墙壁、地板、家具、光线等等——然后才把人物放进去。等她出场时,房间已经占据了画面的大部分,她只能缩小到剩余的空间里。
从紧凑的肖像到整个房间的远景,中间的一切档位,都只是你在这两个极端之间滑动的位置:人物出现得多早,以及她分到多大比例的词汇。
提示: 如果你用 LLM 来写提示词,不必手动搭建这种词序。加上一条指令,比如 "Start by describing the scene around the subject, establishing the room in detail first. Then place the persons inside the room, seen at a distance."(先详细描述主体周围的场景,把房间建立起来,然后再把人物放进房间,从远处看到),LLM 就会自动为你套用"房间优先"的结构。调整这句话,就能把取景滑到你想要的任何位置。
示例:取景阶梯
同一个女孩,同一个房间,四种取景。 一位 22 岁的韩国女性,粉色头发,oversized 连帽衫,一间黑色 RGB 灯光的房间。我唯一移动的东西,是文字。
每条提示词都分为两部分:主体部分和场景部分。往下看时注意两件事:
词汇的天平从主体逐渐倾向场景。
到最远端时,场景排到了主体前面,被最先描述。
这种组合式的转移就是全部的变焦控制。除此之外没有任何别的因素在起作用。
(说明:以下提示词保留英文原文,因为它们是可直接使用的成品提示词——文中所有示例图均由这些提示词原样生成,未经修改。每条下方附中文大意供参考。)
1. 肖像 —— 主体就是整条提示词
A close photo portrait of an 22-year-old korean woman, face to the camera: playful half-smile, winged eyeliner, glossy lips, soft pastel-pink hair, oversized gaming headphones glowing pink on the earcups.
Behind her, just a blur of pink LED light.
中文大意:一张 22 岁韩国女性的近距离肖像照,面对镜头:俏皮的浅笑、飞翘眼线、水润嘴唇、柔和的粉彩色头发、耳罩发着粉光的大号游戏耳机。她身后只有一片模糊的粉色 LED 光。
关键操作:
几乎每一个词都在描述脸和头部
房间只有一个模糊的短句,没有任何细节
大量的面部细节把镜头拉得很近

2. 三分之二身 —— 主体领跑,房间登场
A three-quarter photo shot of an 22-year-old korean woman, framed from just above her head down to her thighs, lounging sideways in a gaming chair. Face to the camera, pastel-pink hair, an oversized hoodie slipping off one shoulder, pink headphones.
Behind her, a black bedroom in soft focus: RGB strip lights and the glow of a monitor.
中文大意:一张 22 岁韩国女性的七分身照片,取景从头顶上方到大腿,侧身慵懒地坐在电竞椅上。面对镜头,粉彩色头发,oversized 连帽衫滑落一侧肩膀,粉色耳机。她身后是一间柔焦的黑色卧室:RGB 灯带和显示器的光。
关键操作:
明确的裁切:"framed from just above her head down to her thighs"(从头顶上方到大腿)
一个姿势(慵懒地坐在椅子上)锚定了身体
房间进入画面,但只有一行柔焦描述

3. 全身 —— 房间在前,然后是完整的人物
A photo of a black gaming bedroom lit by RGB strips: a black-and-pink gaming chair, a desk with a glowing monitor, LED lights tracing the walls, plushies on the bed.
In the chair, full body shot of an 22-year-old korean woman with pastel-pink hair, oversized hoodie and thigh-high socks, smiling cute, her feet touching the carpet.
中文大意:一间被 RGB 灯带照亮的黑色游戏卧室的照片:黑粉配色的电竞椅、放着发光显示器的桌子、沿墙布置的 LED 灯、床上的毛绒玩具。椅子上,一位 22 岁韩国女性的全身照,粉彩色头发,oversized 连帽衫和过膝袜,笑容可爱,双脚踩在地毯上。
关键操作:
房间先于人物被描述
"full body shot"(全身照)直接点明取景
面部只留两个词("smiling cute")
描述脚,就能强迫脚出现在画面里

4. 房间尺度 —— 房间领跑,人物坐在远处
A wide view photo of a whole black gaming room: RGB-washed walls, a desk with dual glowing monitors, shelves of plushies, string lights, a neon sign, and a bed against the back wall.
Sitting on the edge of the bed, an 22-year-old korean woman with pink hair, smiling to the camera in the distance.
中文大意:一整间黑色游戏房的广角照片:被 RGB 光洗染的墙壁、放着双发光显示器的桌子、摆满毛绒玩具的架子、串灯、霓虹灯牌、靠后墙的一张床。坐在床沿上,一位粉发的 22 岁韩国女性,在远处对着镜头微笑。
关键操作:
整个第一段全是房间
人物只有一个短句,并被放置"in the distance"(在远处)
一个放置锚点("on the edge of the bed",坐在床沿)确定了比例——因为床已经被定义在后墙位置

速查表
变焦 = 词汇预算。 画面属于你最先描述、描述最多的东西。
想拉远: 房间在前,人物在后,并且尽量少给脸分配词汇。
想让某个身体部位入镜?直接描述它 ——不要依赖裁切标签。"Standing barefoot on the white rug"(赤脚站在白色地毯上)即使在较紧的取景下也能把脚拉进画面;单靠点名裁切方式往往会失败。
把身体锚定到房间上 ——"feet on the floor"(脚踩在地板上)、"one hand on the desk"(一只手搭在桌上)、"sitting back in the chair"(靠坐在椅子里)。接触点会强迫这些部位入镜,并确定画面比例。
本文中所有示例图均由所示提示词原样生成,未经任何编辑。

