A collaborative approach to image generation

Admin

A collaborative approach to image generation

PASTA uses a two-stage strategy that combines real human feedback with large-scale user simulation to train an AI agent to adapt to individual preferences. The method is designed to address the difficulty of gathering large amounts of real interaction data from users, including privacy concerns.

First, the researchers collected a high-quality foundational dataset with over 7,000 raters’ sequential interactions. Those interactions included prompt expansions generated by a Gemini Flash large multimodal model and corresponding images generated by a Stable Diffusion XL (SDXL) T2I model.

The seed data was then used to train a user simulator that can generate additional data replicating real human choices and preferences. At the heart of the approach is a user model with two parts: a utility model that predicts how much a user will like a set of images, and a choice model that predicts which set they will select when shown several options.

The user model was built using pre-trained CLIP encoders with user-specific components. It was trained with an expectation-maximization algorithm, which lets the system learn individual preferences while also identifying latent “user types” such as clusters of users who tend to prefer animals, scenic views, or abstract art.

The trained user simulator can give feedback on generated images and make selections from sets of proposed images. The researchers said this enabled more than 30,000 simulated interaction trajectories and created a controlled environment for exploring a wide range of user behaviors while training the PASTA agent to collaborate with users.

Source: research.google.

Companies can share verified announcements through Newz9’s international press release submission page.