MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

arXiv:2608.11616 · cs.AI, cs.CV, cs.LG · Submitted 2026-08-13 · Read on arXiv

Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim

KAIST AI

cs.AI, cs.CV, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Project page: https://hchoi256.github.io/projects/mba/ Code: https://github.com/hchoi256/MBA

Code: https://github.com/hchoi256/MBA

Project page: https://hchoi256.github.io/projects/mba

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper introduces MBA-Bench, described as "the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized

Terminology

Summary

The paper introduces MBA-Bench, described as the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. The authors automatically caption images and use GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Evaluation is conducted across six business-oriented criteria using MLLM-as-a-Judge. The paper presents two agents: MBA-b and MBA-k for blind and known evaluation settings, respectively. Both are trained with two novel reward objectives—creativity and feasibility—while MBA-k further optimizes six disclosed criteria. Training uses LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. Results show MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.

The paper argues that existing business ideation approaches remain fundamentally bound to a 'text-in, text-out' paradigm grounded in patents, which is neither realistic nor practical for real-world settings where ideas can be derived from readily available multimodal sources. The authors hypothesize that visual detail is pivotal to distinctive ideation and that images contain unique information that text cannot fully capture. They demonstrate that caption baselines trail across all six business-oriented metrics and that text alone cannot sustain competitive ideation. Multimodal inputs outperform caption-based counterparts by a large margin, even approaching closed-source performance on viability, but still fall short on creativity, due to its heavy reliance on zero-shot MLLMs that tend to produce conventional, homogeneous ideas.

The benchmark comprises 2K images across six domains: General (ADE20K everyday scenes), Spatial Layout (RICO screenshots with densely arranged mobile interface components), Crowding (COCO images with at least ten annotated people), Visual Condition (VisA anomaly images depicting subtle surface defects), Shape & Texture (DTD close-up material surfaces with fine-grained patterns), and Technical Features (DeepPCB circuit-board images with intricate structural details). Images are selected by ranking images on dataset annotations with domain-specific criteria including object diversity, component count and text density, person count and object diversity, a fixed ratio favoring anomalous cases, uniform coverage of texture attributes, and prioritization of defect-annotated images.

Each image is paired with an automatically generated caption and three business questions—cost efficiency, technology, and user experience. The ideation protocol involves three stages: visual query extraction, market evidence retrieval via the DuckDuckGo API, and evidence-augmented ideation, yielding 30K image–caption–question–idea quadruplets. Each idea comprises four fields: title, description, implementation, and differentiation.

Evaluation uses an MLLM-as-a-Judge approach with six business-oriented dimensions from PBIG: Specificity (1–4), Technical Validity (1–4), Innovativeness (1–5), Competitive Advantage (1–4), Need Validity (0–3), and Market Size (0–3). The judge is a strong 78B-parameter MLLM (InternVL2.5-78B) during evaluation, while Qwen2-VL-72B-Instruct is used during training to avoid model bias.

The paper proposes two agents for different deployment settings:

MBA-b (blind) optimizes two task-general objectives: creativity (novelty relative to reference ideas) and feasibility (market relevance and factuality).

MBA-k (known) additionally optimizes the six disclosed evaluation metrics, for eight total.

For feasibility, the paper introduces MBA-Library, an enormous web-sourced resource for feasibility grounding—spanning mobile applications, scientific literature, structured entities, and Wikipedia. This powers two scoring modules: market relevance (using FAISS to retrieve top-k nearest records by cosine similarity) and factuality (following FActScore, where an LLM decomposes ideas into atomic facts verified against Wikipedia passages).

Both agents follow a two-stage training protocol:

  1. SFT: LoRA-based supervised fine-tuning of Qwen2.5-VL-7B-Instruct on question–reference idea pairs, optimizing next-token prediction loss.

  2. GRPO: Group relative policy optimization with setting-specific rewards, where the judge ranks the group along each objective, with each rank converted into a [0, 1] score. Advantages are normalized within groups, and the policy is updated with KL regularization toward the frozen SFT reference.

For MBA-k, GRPO weights are: Specificity 0.12, Technical Validity 0.12, Innovativeness 0.20, Competitive Advantage 0.16, Need Validity 0.12, Market Size 0.08, Creativity 0.10, and Feasibility 0.10. MBA-b uses creativity weight 0.70 and feasibility weight 0.30.

The paper reports that MBA-k achieves the best performance on both Creativity- and Feasibility-related metrics among all models, including closed-source models. Specifically, MBA-k achieves: Specificity 3.99, Technical Validity 3.00, Innovativeness 4.00, Competitive Advantage 3.32, Need Validity 2.94, and Market Size 2.75. MBA-b achieves: Specificity 3.64, Technical Validity 3.12, Innovativeness 3.98, Competitive Advantage 3.19, Need Validity 2.38, and Market Size 2.20.

Statistical analysis using paired two-sided Wilcoxon signed-rank tests shows MBA-k achieves statistically significant improvements in 12 of 18 model–metric comparisons against GPT-5 mini, Gemini 3.1 Pro, and InternVL2.5-8B, with significant improvements in Innovativeness, Competitive Advantage, and Market Size across all comparators.

Domain-wise analysis: MBA-k achieves near-perfect Specificity across all domains. Need Validity and Market Size peak in General and Crowding domains. Technical Validity scores are higher in Technical Features and Visual Condition. Innovativeness and Competitive Advantage achieve high scores in technology-intensive domains and Spatial Layout.

Modality ablation: Caption baselines perform poorly across all domains except General. Multimodal baselines substantially improve performance, especially in image-specific domains. The agents achieve substantial gains in the two creativity-related metrics through GRPO, while SFT alone already improves the four feasibility-related metrics.

Validating training-time rewards: Spearman's ρ = 0.83 for Creativity-Creativity* and ρ = 0.71 for Feasibility-Feasibility*, confirming the two rewards generalize robustly across rubrics oriented toward both original and viable ideation.

Cross-modal mismatch: The paper demonstrates through LPIPS distances that detailed captions fail to preserve fine-grained visual details and even accurate captions cannot reconstruct unverbalizable visual semantics, motivating multimodal business ideation.

MBA-k achieves the lowest semantic failure rate among open-source models at 0.40% (compared to 5.13–12.60% for other open-source baselines) and reduces invalid-format outputs to 3.40% (compared to 16.93–99.80% for other open-source models). All severe semantic failures arise from Technical Validity.

The paper identifies three limitations: (1) MBA focuses on image and text, excluding audio, olfactory, tactile, and other sensory signals; (2) MBA does not model temporal information, though videos can further provide motion, behavioral, and causal context; (3) MBA generates ideas independently of the prospective entrepreneur, ignoring factors like available capital, expertise, location, social network, and risk tolerance. Future work should extend to richer sensory inputs, temporal multimodal reasoning, and personalized business ideation conditioned on user-specific profiles.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

  • Improvement: Integrate a visual feature extraction layer that processes raw images (not just captions) through a vision encoder, then fuses these features with text embeddings before generation. Use domain-specific visual attention mechanisms (e.g., detecting spatial layouts, texture patterns, or defect regions) to condition idea generation.

  • Capability: The system can generate business ideas that leverage unverbalizable visual details—e.g., spotting a crowded retail space from a photo and proposing a queue-management app, or detecting a subtle surface defect in a manufacturing image and suggesting a quality-control service.

  • Improvement: Implement a two-stage training pipeline: (a) supervised fine-tuning on multimodal question–idea pairs, then (b) group relative policy optimization (GRPO) with two reward functions—creativity (novelty against reference ideas via semantic similarity) and feasibility (market relevance via vector retrieval from a web-sourced knowledge library, plus factuality via atomic-fact verification against Wikipedia). Use adaptive reward weighting (e.g., 0.7/0.3 for creativity/feasibility) that can be tuned per deployment context.

  • Capability: The system can produce ideas that are both novel (not generic or repetitive) and grounded in real-world market evidence, reducing hallucinated or impractical suggestions.

  • Improvement: Build a training/evaluation pipeline that categorizes inputs into six visual domains (general scenes, spatial layouts, crowding, visual anomalies, shape/texture, technical features) and applies domain-specific data augmentation (e.g., cropping for spatial layouts, contrast enhancement for defects) and domain-weighted loss functions during fine-tuning.

  • Capability: The system can adapt its ideation strategy to the visual context—e.g., prioritizing technical feasibility for circuit-board images, user-experience ideas for mobile UI screenshots, and market-size considerations for crowded public scenes.

  • Improvement: Replace human or single-model evaluation with a two-model judge system: use a strong 78B-parameter MLLM for final evaluation and a different 72B model during training to avoid self-bias. Implement six-criteria scoring (specificity, technical validity, innovativeness, competitive advantage, need validity, market size) with weighted aggregation, and run paired Wilcoxon signed-rank tests for statistical significance.

  • Capability: The system can self-assess and iteratively improve its outputs against multi-dimensional business criteria, ensuring balanced performance rather than over-optimizing one metric.

  • Improvement: Add a preprocessing module that computes perceptual similarity (e.g., LPIPS) between the input image and its caption to detect information loss. If loss exceeds a threshold, the system automatically switches to full multimodal processing instead of caption-only, and flags the sample for additional visual feature extraction.

  • Capability: The system can reliably handle cases where captions are incomplete or misleading, ensuring it doesn't miss critical visual cues that could differentiate its ideas.

  • Improvement: Implement a structured output validator that enforces the four-field format (title, description, implementation, differentiation) and runs a semantic failure check (e.g., detecting contradictions or invalid technical claims) before final output. Use a fallback mechanism to regenerate with corrected formatting if validation fails.

  • Capability: The system can produce consistently well-structured, error-free business ideas, reducing invalid outputs from 17% (typical open-source) to under 4%, and eliminating semantic failures except in rare technical validity cases.

  • Generate business ideas from any image (e.g., a photo of a cluttered desk, a satellite image, a medical scan) with both creative novelty and market feasibility, outperforming text-only and caption-based systems by 25–77% on business metrics.

  • Adapt to visual domain specifics: It can propose cost-efficiency ideas for manufacturing defect images, technology ideas for circuit boards, and user-experience ideas for mobile app screenshots—without requiring manual domain switching.

  • Self-evaluate and improve: It can score its own outputs on six business criteria, identify weaknesses (e.g., low market size), and retrain or adjust its generation strategy accordingly.

  • Handle incomplete or ambiguous visual input: If an image has complex details that text cannot capture, it will still leverage the raw pixels to generate differentiated ideas, rather than relying on lossy captions.

  • Produce reliable, structured outputs: It can output business ideas in a consistent four-field format with minimal errors, suitable for direct use in entrepreneurial planning or pitch decks.

  • Scale to real-world deployment: With the MBA-Library grounding, it can retrieve current market evidence (apps, papers, entities) to ensure ideas are not only novel but also actionable and relevant to existing markets.

Abstract

Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.

Sources

Related papers