Phota trains a face once to use it with multiple models

Phota Identity V2 creates a visual profile from five photos, then reuses it with multiple models to generate, edit, or restore the same person.

Creating a convincing portrait has become relatively common. Reproducing exactly the same person in another scene, from a new angle, or after several successive edits remains much more difficult. A face becomes elongated, a gaze changes, a jawline shifts, and the character eventually stops looking like the original subject.

Phota Labs is attempting to solve this problem with Identity V2. The service trains a profile from several photographs of a person, then keeps that representation separate from the model responsible for generating the image. The same identity can therefore be used across different engines without having to recreate it for each one.

The approach targets professional portraits, digital doubles, recurring characters, advertising campaigns, and brand storytelling. It can also be used to edit personal photographs, bring several people together in one scene, restore an old image, or render the same subject in different styles.

The process relies on creating a profile. The user uploads several photographs showing the same person from different angles, with varied expressions and lighting conditions. Phota analyzes facial features, hair, skin tone, and certain broader aspects of the person’s appearance, then provides a `profileid`.

This identifier can then be referenced in a prompt. A request such as “create a professional portrait of `[[profileid]]`” is designed to recall the registered person without requiring all their photographs to be uploaded again. Multiple identifiers can also be included in the same prompt to bring different people together.

Phota claims that Identity V2 can be trained from five images in one minute. However, this duration is not stated consistently across its public documentation. The API guide describes a Quick Train mode using five to ten photographs, with training taking around three minutes excluding queue time. Full Train accepts between ten and fifty images and takes approximately eight minutes.

Elsewhere, the general documentation still mentions a delay of up to around twenty minutes. These discrepancies may come from different versions or infrastructures, particularly between the direct API, Phota’s platform, and the integration available through fal. Without a single, up-to-date specification, the one-minute claim should therefore be considered the fastest advertised time rather than a guaranteed duration for every user.

The number of images also affects the expected level of accuracy. Quick Train prioritizes speed and lowers the cost of creating a profile. Full Train remains Phota’s recommended option for production use, with up to fifty photographs providing the most complete representation.

Images that are too similar provide little information about how the subject varies. A profile intended for use across many different scenes benefits from combining front-facing and three-quarter views, multiple expressions, different lighting conditions, and sufficiently sharp framing. Permanent accessories, highly variable hairstyles, or group photographs can also make identification more difficult.

Group photographs remain accepted during training as long as the intended subject is the person who appears most frequently across the submitted set. Phota also supports cats and dogs, although the evaluation published alongside Identity V2 focuses primarily on human faces.

This method differs from the two approaches typically offered by visual generators. The first involves sending reference images with every new request. GPT Image, Nano Banana, Reve, or Seedream must then reinterpret the identity during each generation using only the elements available in their context.

This approach requires no prior training, but its accuracy depends heavily on the photograph selected. A front-facing image may provide insufficient information about a profile view, an extreme expression, or an angle that is absent from the reference. The subject may also drift when the framing, clothing, or scene differs too much from the original image.

The second approach uses a LoRA dedicated to a specific person. This lightweight adaptation learns the subject’s appearance and associates it with a particular model. It can improve resemblance, but generally needs to be adapted or retrained when the user wants to switch engines.

Phota places identity in a third layer. The profile is trained once, then combined with the selected model at request time. Users can therefore take advantage of the strengths of each engine without repeating the entire personalization process.

It is therefore not a new generator designed to replace Nano Banana or GPT Image. Phota operates on top of these models to provide them with knowledge of a specific person. Composition, prompt adherence, text rendering, style, and part of the overall image quality still depend on the selected engine.

The API documentation currently lists Nano Banana 2, used by default, GPT Image 2, two variants of GPT Image 2.5, Qwen Image 2, Flux 2 Dev, Flux 2 Pro, and Seedream 5.0 Pro. Available models and supported resolutions vary depending on the engine.

Identity can be applied to both generation and editing features. Users can place a person in a new environment, change their clothing, expression, position, or camera angle. The Enhance and Remix features are designed respectively to improve a photograph and reproduce the aesthetic of a reference image while attempting to preserve recognized individuals.

Identity V2 claims to handle strong expressions more effectively without systematically pulling the face back toward those found in the training photographs. Phota also emphasizes the preservation of skin tones, skin texture, and distinctive facial features when lighting, pose, or style changes.

These characteristics appear in demonstrations, but not all of them are measured separately in the report. The study primarily evaluates overall facial resemblance, prompt adherence, and visual quality. It does not publish a dedicated test focused solely