This podcast is a live conversation I hosted for Skip Coach, a community of senior product and tech leaders navigating career transitions together. Members get access to live sessions like this one, including real-time Q&A, before the content is shared publicly. If you’re a product or tech leader looking to maximize your career, apply to join Skip Coach — it’s free — to access this content and other exclusive resources.
Listen on YouTube, Spotify, and Apple Podcasts
Brought to you by:
Chargebee—Billing and monetization infrastructure for the AI era
“Wait, I didn’t know the models could do that.” That is how Tulsee Doshi describes the light-bulb moments in her job, and Nano Banana gave her one. Her team was building the image editing capability: hand the model a photo, describe the change you want, and it makes that change. They could see the potential while it was still coming together. Then someone tried it for the first time, and it was magic. “That level of joy as a PM I haven’t experienced before.” Those moments don’t come every day, and they are what “fuel all of the other chaos.”
Tulsee leads product for the Gemini models, the engine that the Gemini app, Search, and every developer on the API build on, plus Google’s video, image, music, and audio models. She has been at Google eleven years, across Search, YouTube, and a stint leading responsible AI, before moving to the model team.
The chaos is the job, and it doesn’t fit any playbook most of us have run. Product management at Google still looks like big tech: planning cycles, launch reviews, hundreds of teams coordinating a release. That machine is alive and well around her. Consumer PM, the kind she practiced at YouTube, is about reading data and building taste. Enterprise PM starts from a customer’s stated problem and a requirements doc. Running product for a frontier model takes a mixture of all three, and then bends each of them. No customer hands the team a requirements document. The team decides what the model should get better at, then has to describe it precisely enough for researchers to build toward, as the researchers can improve anything you can measure. The ground moves with every release, and every team in the company wants something from the next one.
I sat down with Tulsee for the fourth conversation in our Inside PM series, after Meta, Stripe, and Waymo, to understand how product management actually works in that seat, and what a product leader anywhere can take from it.
Five things stood out:
Customer requests get weighed against something the customers can’t see. And the seat that lets you see it has to be established.
A model has a personality, and someone on the product team has to own it. The benchmarks can’t see the part that decides adoption.
“Technical” on this team doesn’t mean computer science. It means being able to say what good looks like.
The best PMs she works with prototype features that don’t work yet. On purpose.
Shipping runs on conviction about a few signals, not a checklist. Each release is a fresh zero-to-one.
Below are the core insights I took from the conversation. The full episode has more.
How a Model Gets Built
The model team is a platform team. Its job is to decide what capable means: for which customers, for which uses, and where the model should shine. Coding, video, long documents, agents that take actions over hours. Those are choices, and they get made before the research starts.
Requests arrive from every direction, and the process for handling them is more concrete than you’d expect. A product team starts by building a prototype on the model as it exists today, because that is the surest way to see where it falls over. The failures get sorted. Some are fixable with a system instruction or better prompting. Some are fundamental to the model and need real research time, and those get turned into evals: a set of examples with a clear definition of what a win and a loss look like. Once a team has one, the researchers have a target. In her words, “nothing can replace a great eval.”
Product managers work at each of those steps. Here is how they run it.
Every Request Gets Weighed Against Where the Research Is Going
The model’s customers come in three kinds: the lab’s own products like Antigravity, its coding tool, other Google teams like YouTube and Gmail, and every developer on the API. Their requirements are disparate, and they need different levels of steadiness. An internal surface can take a live experiment and improve the model incrementally; a developer on the API needs a version that “can’t change under them.” The familiar part of the job is consolidating those asks, and Tulsee’s team does it the way most platform teams do, at the capability level rather than product by product. Deep Research was built into the Gemini app, brought to Search, and turned out to be what enterprise customers in finance and legal were asking for. One capability, three customers.
The less familiar part is that the requirements alone won’t produce an interesting model.
“If you were to simply sort of stack rank all of the asks you were getting in just some list of similarity, you would probably have a boring product on the other side.”
Innovation comes from the researchers pushing the envelope, and the customers asking for features can’t see that work. They “don’t know actually what we’re investing in or where we think the research is going.” So the balancing act is bigger than weighing customer requirements against each other. It is knowing where the research is going, protecting it, and still making sure the model meets a growing number of use cases.
Knowing where the research is going means having a seat with the researchers, and that seat has to be established. Researchers aren’t used to having a product manager in their work, and the lane is new enough that there is no standard role to step into. What Tulsee looks for is what motivates the researcher in front of her. Some are driven by users and want a partner who synthesizes what customers need. Others already have a strong point of view about where the model should go and want help taking that vision “from point A to point B,” not someone announcing they know better. Calibrate to the person and deliver, and “the door opens more and more and more for you to have an opinion earlier and earlier and earlier in that process.” That opinion, on where the research is going, is what the customer requests get weighed against.
A Model Has a Personality, and Somebody Has to Own It
Every new checkpoint arrives with hundreds of benchmark numbers, and none of them makes it easy to say what the model is like to talk to. Tulsee’s word for that gap is vibes, and she treats it as product territory. The failure mode is treating the scores as the product: “you get focused on the numbers and you treat the numbers as the definition of the model.”
The PMs who get the vibes right are the ones who use everything. Simon Tokumine, the PM for NotebookLM, is her example. When Simon says he loves a model, or questions it, she takes it seriously, “because he is someone who deeply uses every model, not just the models that are internal to Google, but also the competitor models.” The instrument is a standing set of prompts, as simple as asking for a joke or as involved as an agentic task, run against each new model and against the competition.
This is the consumer PM’s skill, taste ahead of the data, pointed at a product whose personality changes with each new model.
The Scarce Skill Is Saying What Good Looks Like
There is no shortage of ways to make the model better. Her researchers can hill-climb toward almost anything you point them at, and one of her leads puts it plainly: the hard problem is defining good evals, choosing which of the many possible evals to invest in, and setting the right expectations for what good should be. The scarce skill is saying, precisely, what good looks like.
The practice starts small. “Take 20 prompts,” she says. Twenty examples of the problem, replicable, where a win and a loss are obvious. That forces a vague ask like “make the information architecture better” into something specific enough to align a room on, and it is the seed of an eval the researchers can climb. The best evals in the industry took days and weeks to build, and she calls this “the skill we’re still building and learning too.”
That is also why “technical” means something different on her team, which hires from finance, consulting, and operations as readily as from engineering. “Maybe even more than technical, I mean thorough. Individuals who don’t just look at a number and take that number at face value and say, ah, this is 95%, great. But they actually then go 2 levels lower than 95.” A research PM doesn’t have to be a former researcher. They have to be able to sit next to one, look at the same graph, and hold their own.
Saswat Panigrahi at Waymo said the same thing from the other side of the table: the product leaders who can specify what good looks like are the ones whose careers pay off. It is the one habit here that any PM can practice tomorrow.
Prototype the Features That Don’t Work Yet
“The field changes every week,” Tulsee says. “And therefore with that, your roadmap will change.” Most product teams treat that as a reason to build on solid ground. The best ones she works with do the opposite: “they’re always prototyping something that doesn’t quite work yet.”
The edge has a shape. Not science fiction, but “something where there’s like glimmers of light that like, if something got a little bit better, it would work incredibly well.” NotebookLM is her example. The team builds against today’s latest model, but keeps a set of features the model can’t handle yet. Each new model gets tried against that set. When capability catches up they are first to benefit, and the feedback they send upstream is specific enough to drive the research direction.
The risk runs the other way. “If we try too hard with these models right now to build the perfect product today, you are going to anchor on a product that is actually maybe even too simple for the world of tomorrow.” A rule of thumb from the conversation: if today’s model fully supports the problem, you cut it too shallow; if it doesn’t support it at all, too aggressive. Roughly 60 percent working, 20 percent mediocre, 20 percent reach.
You Can’t Turn Every Launch Light Green
Google’s big-tech launch machine still exists: hundreds of teams, each flipping its light to green before a release goes out. A model team can’t run that way, because each release is a fresh zero-to-one: a new model nobody has used yet. So Tulsee’s team ships on conviction. “A lot of what shipping quickly comes down to is really having conviction in what are the signals that are most important to give you early signal about the model.” Basic evals first, then real users as fast as possible, and two questions: is anything fundamentally broken, and are there glimmers of light? The tell is behavior. The models worth shipping are the ones where “you see that market change in user behavior,” longer conversations, harder tasks, people coming back.
Shipping that way, over and over, is what produces the chaos she talks about. The team is reacting to each new model as it lands, the pace doesn’t let up, and part of her job now is keeping that organized chaos from becoming burnout. “How do we move fast but also move strategically and thoughtfully?” is the question she keeps asking.
What She Would Take From the Lab to Any Product Job
Most of us are not going to work in a lab. The question is what Tulsee would carry with her if she left to run product at a company that builds on models instead of building them. Her answer comes down to four things.
Build first. “I would just focus so much more on... building quickly and prototyping first.” The old cycle was research, requirements, mock, build, test. Now a prototype that gets you 40 percent of the way there speeds up the user research, surfaces the edge cases, and tells you sooner whether the product is real.
Start with twenty prompts. Her advice for anyone trying to get better at evals: twenty examples of the problem, replicable, where a win and a loss are obvious. It is the fastest way to say what you actually want, and when you aren’t the one improving the model, it is how you compare the frontier models against each other and put a number on how well each one does the job.
Keep a few features the model can’t handle yet. Ship what today’s model does well, and prototype the things that almost work. When the next model lands, try them again. The teams that do this move first.
Use every model yourself. Yours and the competitors’. The scores won’t tell you what a model is like to work with, and that is the part your users react to.
None of this needs a lab. It needs a product manager who knows what they want well enough to show it, and who keeps their hands on the models long enough to notice when the answer changes.
Have your own career question? Get personalized guidance at Nikhyl.AI. It’s where the questions keep coming, and where I’ll keep sharing what I’m learning.






