Chat models — NinjaChat

One prompt. Every answer.

The same three prompts, run through every model on this page. Unedited, one take each.

Open a thread

The specimen sheet

Pick a prompt. Read what each model actually said.

Prompt

Write a Python function that merges overlapping intervals, plus three short tests. Keep the whole answer under 40 lines of code.

  • Claude Opus 5

    claude-opus-5
    def merge_intervals(intervals):
        """Merge overlapping intervals.
    
        Args:
            intervals: iterable of (start, end) pairs with start <= end.
    
        Returns:
            New list of disjoint (start, end) tuples sorted by start.
            Touching intervals like (1, 2) and (2, 3) are merged.
        """
        merged = []
        for start, end in sorted(tuple(iv) for iv in intervals):
            if merged and start <= merged[-1][1]:
                prev_start, prev_end = merged[-1]
                merged[-1] = (prev_start, max(prev_end, end))
            else:
                merged.append((start, end))
        return merged
    
    
    def test_overlapping_and_touching():
        assert merge_intervals([(1, 3), (2, 6), (8, 10), (15, 18)]) == [
            (1, 6), (8, 10), (15, 18)]
        assert merge_intervals([(1, 2), (2, 3)]) == [(1, 3)]
    
    
    def test_unsorted_and_nested():
        assert merge_intervals([(5, 7), (1, 10), (2, 3)]) == [(1, 10)]
    
    
    def test_edge_cases():
        assert merge_intervals([]) == []
        assert merge_intervals([(4, 4)]) == [(4, 4)]
        assert merge_intervals([(1, 2), (3, 4)]) == [(1, 2), (3, 4)]
    
    
    if __name__ == "__main__":
        test_overlapping_and_touching()
        test_unsorted_and_nested()
        test_edge_cases()
        print("all tests passed")
    

    A couple of notes on choices worth knowing about:

    • Touching intervals merge. (1, 2) and (2, 3) become (1, 3) because the comparison is start <= prev_end. If your intervals are half-open (e.g. time slots where 2 is an exclusive end), change it to start < merged[-1][1] so they stay separate.
    • Input is not mutated, and the output is a list of tuples regardless of whether you pass lists or tuples in.
    • Invalid intervals aren't validated. A reversed pair like (5, 1) will produce nonsense rather than an error; add a check if the input isn't trusted.

    Complexity is O(n log n) time from the sort, O(n) extra space.

  • GPT-5.6 Sol

    gpt-5.6-sol
    def merge_intervals(intervals):
        """Merge overlapping intervals and return them sorted."""
        if not intervals:
            return []
    
        intervals = sorted(intervals, key=lambda interval: interval[0])
        merged = [list(intervals[0])]
    
        for start, end in intervals[1:]:
            if start <= merged[-1][1]:
                merged[-1][1] = max(merged[-1][1], end)
            else:
                merged.append([start, end])
    
        return [tuple(interval) for interval in merged]
    
    
    assert merge_intervals([(1, 3), (2, 6), (8, 10)]) == [(1, 6), (8, 10)]
    assert merge_intervals([(5, 7), (1, 2), (2, 4)]) == [(1, 4), (5, 7)]
    assert merge_intervals([]) == []
  • GPT-5.5

    gpt-5.5
    def merge_intervals(intervals):
        if not intervals:
            return []
    
        intervals = sorted(intervals)
        merged = [list(intervals[0])]
    
        for start, end in intervals[1:]:
            if start <= merged[-1][1]:
                merged[-1][1] = max(merged[-1][1], end)
            else:
                merged.append([start, end])
    
        return [tuple(interval) for interval in merged]
    
    
    assert merge_intervals([(1, 3), (2, 6), (8, 10)]) == [(1, 6), (8, 10)]
    assert merge_intervals([(5, 7), (1, 2), (2, 4)]) == [(1, 4), (5, 7)]
    assert merge_intervals([]) == []
    
  • GPT-5.6 Luna

    gpt-5.6-luna
    def merge_intervals(intervals):
        """Return overlapping (or touching) intervals merged."""
        if not intervals:
            return []
    
        intervals = sorted(intervals)
        merged = [list(intervals[0])]
    
        for start, end in intervals[1:]:
            if start <= merged[-1][1]:
                merged[-1][1] = max(merged[-1][1], end)
            else:
                merged.append([start, end])
    
        return [tuple(interval) for interval in merged]
    
    
    # Tests
    assert merge_intervals([]) == []
    assert merge_intervals([(1, 3), (2, 6), (8, 10)]) == [(1, 6), (8, 10)]
    assert merge_intervals([(5, 7), (1, 2), (2, 4)]) == [(1, 4), (5, 7)]
    
  • Gemini 3.7 Flash

    gemini-3.7-flash
    def merge_intervals(intervals):
        if not intervals:
            return []
        # Sort intervals based on the start time
        intervals.sort(key=lambda x: x[0])
        merged = [intervals[0]]
    
        for current in intervals[1:]:
            prev = merged[-1]
            # Overlap exists if current start <= previous end
            if current[0] <= prev[1]:
                prev[1] = max(prev[1], current[1])
            else:
                merged.append(current)
    
        return merged
    
    
    # Tests
    assert merge_intervals([[1, 3], [2, 6], [8, 10], [15, 18]]) == [
        [1, 6],
        [8, 10],
        [15, 18],
    ]
    assert merge_intervals([[1, 4], [2, 3]]) == [[1, 4]]
    assert merge_intervals([[1, 2], [3, 4]]) == [[1, 2], [3, 4]]
    
    print("All tests passed!")
    
  • Kimi K3

    kimi-k3
    def merge_intervals(intervals):
        """Merge overlapping intervals. Each interval is [start, end]."""
        merged = []
        for start, end in sorted(intervals):
            if merged and start <= merged[-1][1]:
                merged[-1][1] = max(merged[-1][1], end)
            else:
                merged.append([start, end])
        return merged
    
    
    def test_merge_intervals():
        # Overlapping and disjoint intervals
        assert merge_intervals([[1, 3], [2, 6], [8, 10], [15, 18]]) == [[1, 6], [8, 10], [15, 18]]
        # Touching intervals merge; unsorted input handled
        assert merge_intervals([[5, 7], [1, 4], [4, 5]]) == [[1, 7]]
        # Empty and single-interval edge cases
        assert merge_intervals([]) == [] and merge_intervals([[2, 3]]) == [[2, 3]]
    
    
    test_merge_intervals()
    print("All tests passed.")
    

    The function sorts intervals by start time, then walks through once, extending the last merged interval when there's overlap. Runs in O(n log n) time.

  • GLM 5.2

    glm-5.2
    def merge_intervals(intervals):
        if not intervals:
            return []
        intervals.sort(key=lambda x: x[0])
        merged = [intervals[0]]
        for start, end in intervals[1:]:
            last_end = merged[-1][1]
            if start <= last_end:
                merged[-1][1] = max(last_end, end)
            else:
                merged.append([start, end])
        return merged
    
    # Three short tests
    assert merge_intervals([[1,3],[2,6],[8,10],[15,18]]) == [[1,6],[8,10],[15,18]]
    assert merge_intervals([[1,4],[4,5]]) == [[1,5]]
    assert merge_intervals([[1,2],[3,4]]) == [[1,2],[3,4]]
    print("All tests passed")
    
  • GLM 4.7

    glm-4.7
    def merge_intervals(intervals):
        if not intervals:
            return []
        intervals.sort(key=lambda x: x[0])
        merged = [intervals[0]]
        for current in intervals[1:]:
            last = merged[-1]
            if current[0] <= last[1]:
                last[1] = max(last[1], current[1])
            else:
                merged.append(current)
        return merged
    
    # Tests
    assert merge_intervals([[1,3],[2,6],[8,10],[15,18]]) == [[1,6],[8,10],[15,18]]
    assert merge_intervals([[1,4],[4,5]]) == [[1,5]]
    assert merge_intervals([]) == []
    
  • MiniMax M3

    minimax-m3
    def merge_intervals(intervals):
        if not intervals:
            return []
        intervals = sorted(intervals, key=lambda x: x[0])
        merged = [intervals[0][:]
        for start, end in intervals[1:]:
            if start <= merged[-1][1]:
                merged[-1][1] = max(merged[-1][1], end)
            else:
                merged.append([start, end])
        return merged
    
    
    # Tests
    assert merge_intervals([[1, 3], [2, 6], [8, 10], [15, 18]]) == [[1, 6], [8, 10], [15, 18]]
    assert merge_intervals([[1, 4], [4, 5]]) == [[1, 5]]
    assert merge_intervals([]) == []
    assert merge_intervals([[1, 4], [2, 3], [5, 7]]) == [[1, 4], [5, 7]]
    print("All tests passed!")
    
  • Qwen 3.8 Flash

    qwen-3.8-flash
    def merge_intervals(intervals):
        intervals = sorted(intervals)
        merged = []
        for start, end in intervals:
            if merged and start <= merged[-1][1]:
                merged[-1][1] = max(merged[-1][1], end)
            else:
                merged.append([start, end])
        return merged
    
    assert merge_intervals([[1, 3], [2, 6], [8, 10]]) == [[1, 6], [8, 10]]
    assert merge_intervals([[1, 4], [4, 5]]) == [[1, 5]]
    assert merge_intervals([[2, 3], [4, 5], [6, 7]]) == [[2, 3], [4, 5], [6, 7]]
    
  • GPT-OSS 120B

    gpt-oss-120b
    def merge_intervals(intervals):
        """Merge overlapping intervals.
    
        Args:
            intervals: List of [start, end] pairs.
    
        Returns:
            List of merged intervals sorted by start.
        """
        if not intervals:
            return []
        intervals = sorted(intervals, key=lambda x: x[0])
        merged = [intervals[0]]
        for cur in intervals[1:]:
            prev = merged[-1]
            if cur[0] <= prev[1]:               # overlap
                merged[-1] = [prev[0], max(prev[1], cur[1])]
            else:
                merged.append(cur)
        return merged
    
    
    # --- short tests -------------------------------------------------
    assert merge_intervals([[1, 3], [2, 6], [8, 10], [15, 18]]) == [[1, 6], [8, 10], [15, 18]]
    assert merge_intervals([[1, 4], [4, 5]]) == [[1, 5]]
    assert merge_intervals([]) == []
    
    print("All tests passed.")
    
Prompt

In about 120 words, tell a first-time founder when to choose a monolith over microservices, and the one signal that means it's time to split.

  • Claude Opus 5

    claude-opus-5

    Start with a monolith. Almost always.

    If you're pre-product-market-fit, have fewer than roughly 15 engineers, and don't yet know which parts of your product will matter, a single deployable app is the right call. Microservices buy you independent scaling and team autonomy, but they charge you in distributed tracing, network failures, versioned APIs, duplicated auth, and a deploy pipeline that eats a founding engineer's month. You'll pay that tax daily while your actual risk is that nobody wants the product. Keep the monolith modular — clear internal boundaries, no cross-module database reads — so splitting later is refactoring, not rewriting.

    The signal: teams start blocking each other on deploys. When shipping a small change means coordinating with two other groups, or one component's scaling needs are distorting the whole system, extract that piece — and only that piece.

  • GPT-5.6 Sol

    gpt-5.6-sol

    For a first-time founder, choose a monolith by default. It is faster to build, easier to test and deploy, cheaper to operate, and simpler to debug while your product, traffic patterns, and team boundaries are still changing. Keep it modular: define clear domain boundaries, avoid tangled dependencies, and give each module ownership of its data and interfaces. Do not adopt microservices merely to appear scalable; most startups need product learning more than infrastructure flexibility.

    The signal to split is repeated, measurable pressure for one well-defined component to operate independently—because its deployment frequency, scaling needs, reliability requirements, or ownership consistently slows the rest of the application. When that constraint keeps recurring despite good modular design, extract that bounded component first, not the entire monolith.

  • GPT-5.5

    gpt-5.5

    Choose a monolith by default as a first-time founder.

    If your product is still changing, your team is small, and you do not yet know which parts of the system will matter most, a monolith is usually faster, cheaper, and safer. You get one codebase, one deployment path, simpler debugging, fewer infrastructure decisions, and less coordination overhead. Most early startups need learning speed more than architectural purity.

    Make it a modular monolith: keep clean boundaries between domains like billing, users, notifications, and analytics, but deploy them together.

    The signal it is time to split is this:

    One part of the system has a clearly different scaling or deployment need, and keeping it inside the monolith is repeatedly slowing you down.

    Until then, stay monolithic.

  • GPT-5.6 Luna

    gpt-5.6-luna

    As a first-time founder, start with a modular monolith unless you have a proven reason not to. It is faster to build, test, deploy, debug, and change when you have one codebase, database, and simple operational model. Keep boundaries clean—separate modules, APIs, and ownership—so you preserve the option to extract services later. Choose microservices initially only when a domain has genuinely different scaling, reliability, security, or delivery needs, and you can afford the operational cost and team complexity.

    The signal to split is repeated deployment or scaling friction: one component needs to be released or scaled independently often enough that the monolith slows delivery or wastes resources. Extract that boundary first, measure the result, and avoid distributing everything by default.

  • Gemini 3.7 Flash

    gemini-3.7-flash

    Always start with a monolith. As a first-time founder, your primary risk is running out of money before finding product-market fit, not handling millions of users. A monolith lets a small team move fast, refactor effortlessly, and deploy without the crippling overhead of distributed systems, complex DevOps, and network debugging.

    The single signal that it’s time to split is organizational friction, not traffic volume.

    When multiple engineering teams are constantly stepping on each other’s toes—causing deployment gridlock, painful merge conflicts, and blurred system ownership—it is time to extract microservices. Until your team's size breaks the development process, keep everything in one codebase and focus entirely on shipping value to your customers.

  • Kimi K3

    kimi-k3

    Choose a monolith while your team is small and your product is still finding its shape. Early on, iteration speed beats everything, and microservices tax you with deployment pipelines, distributed debugging, and consistency headaches you haven't earned. Worse, you don't know your domain boundaries yet—split wrong and you get a distributed monolith: all the pain, none of the upside.

    The one signal to split: team coordination pain. When engineers routinely block each other's deploys, trip over the same code, and clear ownership boundaries have emerged, extract services along those lines. Split for organizational reasons, not because of a performance chart—most scaling problems have cheaper answers.

  • GLM 5.2

    glm-5.2

    As a first-time founder, your primary goal is finding product-market fit, not building scalable infrastructure. Choose a monolith. Microservices introduce massive operational overhead—complex deployment pipelines, network latency, and distributed system debugging—that a small team simply cannot afford. A monolith lets you move fast, iterate quickly, and keep your entire codebase in one easily understood place.

    When is it finally time to split? The one undeniable signal is a deployment bottleneck caused by team growth. If you have multiple engineers working on different features, but they are constantly stepping on each other's toes, breaking unrelated code, or waiting in a deployment queue, it’s time to decouple. Until then, stay monolithic.

  • GLM 4.7

    glm-4.7

    Stick with a monolith initially. It offers the fastest path to launch, simplifies debugging, and avoids the operational "tax" of distributed systems. As a first-time founder, your resources are finite; you need to focus entirely on finding product-market fit rather than managing complex infrastructure. A monolith allows your small team to move quickly without worrying about network latency or eventual consistency.

    The definitive signal to split is organizational friction. It is time to migrate to microservices when your team size makes coordination painful. If merging code causes constant conflicts, or if deploying a single feature is blocked by unrelated components, your architecture has become a bottleneck. Split when your code structure starts slowing down your people, not your servers.

  • MiniMax M3

    minimax-m3

    Start with a monolith. For a first-time founder, this is almost always the right call. Build one well-structured app, ship it fast, and let real users—not architectural diagrams—shape what you actually need. Microservices solve coordination problems you don't have yet: team boundaries, independent deploy pipelines, divergent scaling profiles. Premature splitting burns weeks on infrastructure, debugging network failures, and managing data consistency across services—all while your product is still searching for fit.

    The one signal that means it's time to split: a single component is causing deploy pain for the whole team. If one module (say, payments or search) needs to scale or change on its own cadence and is forcing every other team to coordinate releases, extract it. Until that friction shows up, keep everything in one deployable unit. Speed and clarity beat elegance at this stage.

  • Qwen 3.8 Flash

    qwen-3.8-flash

    As a first-time founder, choose a monolith first. Your biggest risk is building the wrong thing, not scaling the right thing. A monolith lets you move fast, share one database, deploy together, and change your mind without distributed-system headaches. It keeps operations simple: one codebase, one CI/CD pipeline, one mental model. Microservices add latency, versioning, observability, and coordination costs that rarely pay off before product-market fit.

    The one signal to split: when a service has its own stable, high-demand domain that repeatedly blocks other teams—different scaling needs, release cadence, or ownership—then extract it. If your bottleneck is “too many changes conflict,” not “users are overwhelming the system,” fix modularity first. Split only to remove a real constraint.

  • GPT-OSS 120B

    gpt-oss-120b

    Choose a monolith when you’re still figuring out product‑market fit. A single codebase lets you ship features fast, keep deployment simple, and avoid the operational overhead of service discovery, monitoring, and version coordination that micro‑services demand. It’s especially wise when your team is small (1‑5 engineers), the domain is not yet stable, and you need to iterate on UI, business rules, and data model without negotiating contracts across services.

    Signal to split: the moment a single team repeatedly hits a “cross‑team bottleneck” because two unrelated parts of the app must be changed together—e.g., a database schema change for the billing module forces a simultaneous rewrite of the user‑profile code. That contention indicates the codebase has grown enough that independent deployment and scaling would relieve the friction, and it’s time to break out micro‑services.

Prompt

Rewrite this sentence three ways, each in a clearly different tone (plain, playful, executive): "Our app makes budgeting easy." Label each.

  • Claude Opus 5

    claude-opus-5

    Plain: Our app helps you manage your budget without the hassle.

    Playful: Budgeting used to be a chore — our app turns it into a two-tap victory lap.

    Executive: Our platform streamlines personal budgeting, giving users clear visibility and control over their spending.

  • GPT-5.6 Sol

    gpt-5.6-sol

    Plain: Our app simplifies budgeting.

    Playful: Budgeting just got a whole lot easier—and more fun.

    Executive: Our app streamlines budgeting for faster, more confident financial decisions.

  • GPT-5.5

    gpt-5.5

    Plain: Our app helps you manage your budget easily.

    Playful: Budgeting? Our app makes it a breeze.

    Executive: Our app simplifies budget management for greater financial control.

  • GPT-5.6 Luna

    gpt-5.6-luna

    Plain: Our app simplifies budgeting.

    Playful: Budgeting just got a whole lot less boring.

    Executive: Our app streamlines budgeting to improve financial management.

  • Gemini 3.7 Flash

    gemini-3.7-flash

    Plain: "We help you track and manage your money without the hassle."

    Playful: "Budgeting without the boring spreadsheets or math headaches? Yep, we made it that simple!"

    Executive: "Our platform streamlines personal financial governance, empowering users with intuitive oversight and strategic control over their capital."

  • Kimi K3

    kimi-k3

    Plain: Our app simplifies budgeting.

    Playful: Budgeting doesn't have to be a chore—our app makes it a total breeze! 🎉

    Executive: Our solution streamlines budget management, enabling users to achieve financial clarity with minimal effort.

  • GLM 5.2

    glm-5.2

    Plain: Our app helps you manage your money simply and without hassle.

    Playful: Say goodbye to budget headaches—our app makes tracking your pennies a piece of cake!

    Executive: Our platform streamlines financial planning and optimizes resource allocation.

  • GLM 4.7

    glm-4.7

    Plain Budgeting is simple with our app.

    Playful Say goodbye to money stress and hello to budgeting bliss!

    Executive Our platform streamlines financial management to ensure optimal efficiency.

  • MiniMax M3

    minimax-m3

    Plain: "Our app helps you manage your budget."

    Playful: "Budgeting boring? Not anymore — our app makes counting cash actually fun!"

    Executive: "Our application delivers a streamlined, user-centric solution that simplifies personal budgeting at scale."

  • Qwen 3.8 Flash

    qwen-3.8-flash

    Plain: Our app simplifies budgeting.

    Playful: Budgeting? Consider it handled—with zero fuss and maybe a little fun.

    Executive: Our app streamlines the budgeting process, enabling users to manage finances with clarity and efficiency.

  • GPT-OSS 120B

    gpt-oss-120b

    Plain:
    Our app simplifies budgeting.

    Playful:
    Turn the hassle of budgeting into a breeze—our app does the heavy lifting for you!

    Executive:
    Our platform delivers streamlined, user‑centric budgeting solutions that empower efficient financial management.

September 2, 2026Full run with latency and cost →

The roster

Questions

Are these AI models free to use?
Yes, free to start. A free account can chat with the models on this page; plans add more messages and the flagship models at full length.
Which AI model is the best in 2026?
There is no single best. In our September shootout Opus 5 gave the most thorough code and the strongest writing, GPT-5.6 Terra was fastest to the first token, and gpt-oss-120b and GPT-5.6 Luna were the cheapest per answer. Read the answers above and pick the one that fits your work.
What is the difference between GPT-5.6 Sol, Terra and Luna?
Three sizes of one family. Sol is the deep-reasoning flagship, Terra the balanced tier for coding and agents, Luna the fast, inexpensive tier. We measured all three on the same prompts, side by side on this page.
Can I use several models in one conversation?
Yes. Switch models mid-thread, or compare up to four on the same prompt. Your history stays in one place regardless of the model.
What does context window mean?
How much text the model can consider at once, in tokens. A million tokens is roughly a long novel or a mid-sized codebase, so models with that window can read a whole repository in one go.
Which models are open weights?
Kimi K3, GLM-5.2, GLM-4.7, MiniMax M3, Qwen 3.8 Flash, gpt-oss-120b, DeepSeek and Llama 4 publish their weights. They run on NinjaChat like any other model, and you can self-host the same weights elsewhere.

Ask it something hard.

Start a chat