NVIDIA Nemotron 3.5 Lightning and OpenHands: A Faster Path to Always-On Software Engineering Agents

Written by
Olivia Greene
Published on
AI coding agents are moving from one-off prompts to always-on systems. The first wave proved that agents can write code, run commands, inspect repositories, fix bugs, and help developers move faster. But the next wave is bigger than a better chat interface or a smarter autocomplete. The next wave is about agents that run in the background, respond to events, complete multi-step workflows, and operate across teams, repositories, and environments.
That shift changes what matters. For a developer using an agent locally, the question is usually: Can this agent help me finish the task in front of me? For a platform team, the question becomes: How do we run many agents safely, repeatedly, and cost-effectively across the organization? And for an enterprise, the question is: How do we stay in control of the models, data, execution environments, and workflows behind those agents?
That is why OpenHands is supporting the launch of NVIDIA Nemotron 3.5 Lightning, a customizable open model that gives control over always-on agents.
Nemotron 3.5 Lightning points to where agentic software development is going. Not one model for every task, but systems that route the right work to the right model, in the right environment, with the right controls.
OpenHands gives teams the open-source, model-agnostic platform to run those agents from local development to enterprise-scale workflows.
Always-on agents need a system of models
A coding agent does not just answer a prompt. It gathers context, reads files, searches the codebase, plans, edits, runs tests, reasons over errors, and decides what to try next. Sometimes it asks for human input. Sometimes it opens a pull request. Sometimes it runs again when a new issue, Slack message, CI failure, or scheduled task appears.
Every step calls a model but not every step needs the same model. Some tasks need frontier-level reasoning. Others are repetitive, high-volume, domain-specific, or latency-sensitive.
-
A security remediation workflow might need one model to reason through the root cause and another to classify findings at scale.
-
A PR review automation might need one model for deep code analysis and another for structured summaries.
-
A developer working locally may want a fast model that can run close to their code, while an enterprise platform team may need the flexibility to deploy in a controlled environment.
The future of agent infrastructure is not a single-model stack. It’s a model-agnostic system where teams can choose the right model for each task based on performance, cost, latency, data policy, and deployment requirements.
That is the world OpenHands is built for.
What NVIDIA Nemotron 3.5 Lightning brings to agentic workflows
NVIDIA Nemotron 3.5 Lightning is a customizable, open 30B MoE model with 3B active parameters, distilled from NVIDIA’s frontier Nemotron 3 Ultra model. It is designed for always-on agents and built to run across local systems, edge environments, data centers, and the cloud.
The key point is not just that it is fast. It is that it is fast in a way that matters for agents.
Agents do not make one model call and stop. They make many calls across long-running workflows. Token generation speed, throughput, and deployment flexibility directly affect how quickly agents can complete specialized tasks, how many workflows can run in parallel, and how practical it becomes to run agents continuously.
Preliminary performance reported by NVIDIA showed Nemotron 3.5 Lightning delivers up to 4x higher throughput compared with Gemma and 1.7x higher throughput compared with Qwen in NVIDIA’s H100 configuration. On DGX Spark, NVIDIA reports 96.3 TPS/user, compared with 51.0 TPS/user for Qwen v3.6 35B A3B in the same test setup.
For agent builders, that speed matters because agent workflows are step-heavy by nature. Faster token rollout means agents can move through more steps, more quickly, especially for specialized tasks that happen repeatedly across a team or organization.
Nemotron 3.5 Lightning also gives teams more control over how agents are powered. It is open, customizable, post-trainable for specialized workflows, and deployable across NVIDIA infrastructure, including DGX Spark, DGX Station, RTX PRO, RTX, H100, H200, A100, L40S, GB200/B200, and GB300/B300.
This is exactly the kind of model progress that makes always-on agents more practical.
OpenHands gives teams the platform to run agents their way
A model alone does not make an agent useful in production. The agent stack is starting to look more like real infrastructure: models, harnesses, runtimes, sandboxes, integrations, policies, audit trails, cost controls, and deployment targets.
That is the role OpenHands plays. OpenHands is the open source platform for building and running AI coding agents. Developers can start locally with Agent Canvas, run agents against real repositories, connect the tools they already use, and experiment with automations. As those workflows become valuable, teams can move them into shared environments where agents keep running in the cloud, respond to GitHub, Slack, and scheduled triggers, and become repeatable team workflows. At enterprise scale, organizations can add governance, visibility, access control, cost management, audit logs, and self-hosted deployment.
With Nemotron 3.5 Lightning, OpenHands users get another powerful model option for specialized agent workflows. That matters for three reasons:
First, developers get more choice. OpenHands is model-agnostic by design so teams can use the models that fit their cost, performance, and policy requirements without being tied to a single provider.
Second, platform teams get more control. Nemotron 3.5 Lightning is designed to be customized and deployed across environments. OpenHands gives teams the control layer to decide where agents run, what they can access, what actions they can take, and how their activity is tracked.
Third, enterprises get a path from experimentation to production. A developer can start with local agent workflows. A team can turn those workflows into shared automations. An organization can run agents across repositories and teams with the infrastructure required for governance, visibility, and repeatability.
That continuity is important. The journey from local to production should not require switching tools, rewriting workflows, or locking into a single model provider.
Built for high-volume engineering workflows
Nemotron 3.5 Lightning is especially relevant for high-volume, specialized agent workflows.
In OpenHands, those workflows can look like:
-
Reviewing pull requests before a human starts their day.
-
Investigating failed CI runs and suggesting fixes.
-
Updating documentation when code changes.
-
Remediating dependency issues across many repositories.
-
Triaging issues from GitHub, Slack, Linear, Jira, or other systems.
-
Preparing structured summaries for engineering managers.
-
Applying repetitive migrations or refactors with human review.
-
Running security checks and preparing findings for analysts.
These are not one-off prompts. They are repeatable workflows. They are also the kinds of tasks where model choice should be flexible.
Some workflows may need a frontier model. Some may need a small, fast, specialized model. Some may need to run locally. Some may need to run in a self-hosted enterprise environment. Some may need to be post-trained for a company’s codebase, policies, or domain.
OpenHands gives teams the orchestration layer to make those choices without locking the whole agent system to one model or one deployment pattern.
The next agent stack is open, fast, and model-agnostic
OpenHands was built on a simple belief: teams should be able to inspect, extend, and control the systems that run their agents.
That is why open source is not a side note for us. It is the foundation. The OpenHands core is open source, with more than 80,000 GitHub stars, millions of downloads, and hundreds of contributors. Developers can see how agents work, adapt the platform to their own systems, and build workflows that fit how their teams actually ship software.
Nemotron 3.5 Lightning fits that philosophy.
A customizable open model gives teams more control over model behavior, data handling, deployment, and specialization. A model-agnostic agent platform gives teams more control over how that model is used in real workflows. Together, they move agentic software development away from closed, single-path tools and toward flexible infrastructure that teams can actually operate.
The future of software engineering agents will not be defined by a single model, a single interface, or a single vendor. It will be defined by open, composable systems that let teams run agents on their terms.
OpenHands is proud to support NVIDIA Nemotron 3.5 Lightning and help bring fast, customizable open models into the agent workflows where real software engineering work gets done.
Coding agents are no longer just something you prompt. They are becoming systems you run. OpenHands gives you the control center to run them your way.
About OpenHands
OpenHands is the open-source platform for building and running AI coding agents, with the interface, automations, and control layer needed to go from a single local agent to a system running across an entire organization. The mission is to make agent-based software development accessible, transparent, and controllable by default. That starts in the open. The core framework is open source, giving developers and platform teams full visibility into how agents execute work and interact with their systems.
Developers can start locally with Agent Canvas, connect the models and tools they already use, and run agents against real repositories. As workflows mature, teams can turn one-off agent wins into repeatable automations that run on schedules, respond to GitHub and Slack events, review PRs, fix CI failures, update docs, and keep working when a laptop is closed. For organizations scaling agent usage, OpenHands adds the control layer required to run agents safely across teams, repositories, and environments.
Join our community to collaborate, share use cases, and contribute to the project as it evolves. Download OpenHands to start running agents locally, automate real engineering workflows, and scale them when you’re ready.
Get useful insights in our blog
Insights and updates from the OpenHands team
Sign up for our newsletter for updates, events, and community insights.
