Product Analyst
Product, IT
Bengaluru, Karnataka, India
Posted on Jul 28, 2026
Product Analyst
Bengaluru, India (Onsite)
Product
In office
Full-time
Reports to: Head of Product
Team: Product - a horizontal role supporting whichever pod needs rigorous analysis in a given week (Observability, Workflows, Graph Agents, Infra), with work assigned and prioritized weekly rather than owned by a single pod
About Bolna
Bolna is a YC-backed voice AI orchestration platform built for the Indian market - powering multilingual, vernacular voice agents (Hindi, Hinglish, Tamil, and 10+ languages) at sub-500ms latency across collections, recruitment, sales, and e-commerce use cases. We're an orchestration layer, not a model company: our moat is outcome-labeled vernacular data, a rigorous evaluation infrastructure, and a growing taxonomy of how Indian enterprise voice AI actually fails in production.
Why this role exists
Product decisions at Bolna increasingly hinge on rigorous, code-mixed-aware data analysis - and not just one kind. On one side there's model/eval rigor: LLM benchmarking for post-call intelligence, ASR/WER evaluation, inter-rater reliability on human-labeled calls, routing/latency economics. On the other there's product and growth insight: understanding where self-serve users drop off in their journey, what patterns emerge across lakhs of monthly calls, which use cases and configs are actually working. Both currently sit with the Head of Product alongside strategy and roadmap ownership. We need a dedicated analyst to own the execution and the recurring cadence across both - freeing product leadership to act on findings rather than produce them.
What you'll do
Model & evaluation analysis
- LLM & model benchmarking: Run structured comparisons across model providers (e.g. Sarvam, DeepSeek, Gemini, Claude variants) for tasks like post-call extraction and LLM-as-judge scoring - evaluating cost, accuracy, fill rate, and TTR, with particular attention to Hinglish and code-mixed content.
- Evaluation infrastructure: Build and maintain LLM-as-judge pipelines (e.g. using DeepEval), design and track eval metrics, and run inter-rater reliability analysis (e.g. Krippendorff's alpha) across human call reviewers.
- Golden dataset creation: Support the construction of golden datasets for ASR/transcript labeling - including flagging conventions like code-switch scripting (Devanagari vs. Roman) and transliteration normalization for decision before scoring is meaningful.
- ASR/voice benchmarking: Evaluate WER and related quality metrics across ASR providers and models for Indic languages; work with public benchmarks and academic references where relevant.
- Infra & latency analytics: Analyze routing, latency, and cost data (e.g. Azure PTU utilization, percentile latency distributions) to inform infrastructure and routing decisions.
- Agent behavior analytics: Support population-level analysis of graph agent behavior - node-level aggregates, designed-vs-observed graph diffs, stuck-in-loop detection and similar failure-pattern metrics.
Product & growth insight generation
- Self-serve journey analysis: Instrument and analyze the self-serve funnel (signup → activation → habit → expansion), identifying where users drop off and surfacing friction points for the product team to act on.
- Cross-customer call insights: Mine aggregate call data across customers and use cases for patterns - completion rates by use case, which model/config combinations perform best where, emerging failure patterns across prompt templates - feeding both internal roadmap decisions and future customer-facing benchmarking.
- Ad hoc product analysis: Be the fast, reliable "pull me the data on X" resource for pod PMs - usage patterns, cohort behavior, feature adoption - turning raw usage data into a clear, actionable read.
Across both
- Reporting & tooling: Build repeatable dashboards/scripts (not one-off notebooks) so these analyses run as an ongoing cadence, not a one-time project; present findings to product, ML, and infra stakeholders in a form they can act on.
What we're looking for
Must-have
- 0–2 years of experience - new grad or early-career, in a data/product analyst, applied ML, or research-adjacent role. We're hiring for raw analytical strength and trainability, not a finished track record.
- Strong SQL and Python (pandas at minimum) from coursework, internships, or prior work - comfortable writing and debugging your own queries and scripts without hand-holding.
- Solid statistical fundamentals (distributions, basic hypothesis testing, agreement/reliability concepts) - you don't need production experience with inter-rater reliability metrics, but you should pick up a new stats concept quickly when pointed at it.
- Genuine comfort with ambiguous, messy real-world data - able to notice when something looks off and flag it clearly, even if the resolution isn't yours to call.
- Basic funnel/cohort analysis instincts - comfortable thinking in terms of drop-off stages and segments, even if you haven't formally run a growth funnel before.
- Native or near-native fluency in Hindi (or another Indian language) and English, with comfort reading/labeling code-mixed (Hinglish) text.
Strong plus
- Any exposure to LLM evaluation concepts (prompt-based scoring, LLM-as-judge), speech/ASR evaluation (WER, transcript QA), or eval frameworks like DeepEval - even coursework or personal projects count.
- Any exposure to product/growth analytics - funnel analysis, retention curves, cohort behavior - from a prior role, internship, or personal project.
- Familiarity with cloud inference economics (e.g. token-based billing, provisioned throughput models) is a bonus, not expected.
- A portfolio of self-directed analysis (a project, competition, or writeup) that shows you go looking for the "so what," not just the number.
First 90 days - success looks like
- Running the LLM benchmarking comparison for post-call extraction end-to-end under direction — executing the sweep, producing clean cost/accuracy/fill-rate tables - with judgment calls escalated rather than made solo.
- Contributing meaningfully to the golden dataset build: flagging inconsistencies in transliteration/script conventions clearly enough that a fast decision can be made, then applying that decision consistently across the dataset.
- Running the inter-rater reliability pipeline on a recurring basis once it's set up, without needing to redesign it each time.
- Has produced a clear, first-pass view of the self-serve funnel (signup → activation → habit → expansion) with at least one concrete drop-off point identified and flagged for the team to act on.
- Has become the reliable first pass on "pull me the data on X" for at least two of: model benchmarking, ASR evaluation, routing/latency, agent behavior analytics, self-serve/usage patterns - freeing senior time for interpretation rather than execution.
What we offer
- Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.
- Growth paths: Joining Bolna means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.
- Learning & development: Bolna proactively supports professional development, the processes for which are being set up.
- Competitive compensation + meaningful ESOP
- Bengaluru office, in-person team collaboration
Req ID: R44