lunar
Let’s talk
Lunar · Enterprise AI Lab

RL environments.
Small language models.

We build reinforcement learning environments and train compact language models for your business tasks. From experiment design to deployment, with your data and your team.

An AI lab dedicated to your context, from environment to model.

  • Custom environments and rewards
  • Models evaluated on your tasks
  • Deployment in your infrastructure
From context to trainingExample projects

An environment for learning to decide.

The challengeLearn to allocate requests across teams within capacity and deadline constraints.

Episode 01 / 04Illustrative simulation
  1. Observe
  2. Act
  3. Reward
  4. Learn
RL environment
Request #01
Agent
Team AFull
Capacity3 / 3
Team BAvailable
Capacity1 / 3
Action values
A0.00B0.00

Simplified example: +1 for allocating with capacity; −1 for choosing a full team.

Scope and success criteria defined with your team.

Two specialties. From applied research to production.

RL environmentsSmall Language ModelsTask-level evaluationPrivate infrastructure

Your training environment. Your specialized model.

RL environments and Small Language Model (SLM) projects, delivered independently or combined to fit the challenge. Each workstream has defined deliverables and evaluation criteria.

RL environments

The right context for learning to decide.

We turn your business rules, data, and interactions into reinforcement learning environments. We define observations, actions, rewards, and scenarios to train and evaluate agents or policies.

  • Versioned environments and training scenarios
  • Task-specific rewards and constraints
  • Policy evaluation against a baseline
Small Language Models

Compact models. Well-defined tasks.

We select and specialize smaller language models for your domain. Through curation, fine-tuning, and distillation, we target the right balance of quality, cost, and latency for your tasks.

  • Curated datasets and separate test sets
  • Model fine-tuning and distillation
  • Evaluation and serving on the selected hardware

From dataset to deployment, with your team.

Training data

Curation of examples, trajectories, and quality criteria for each experiment.

Reproducible evaluations

Baselines, test scenarios, and reports to decide what goes into production.

Infrastructure and serving

Training and inference pipelines sized for your environment.

Operational integration

Models and policies connected to agents, tools, and internal systems.

From hypothesis to a model in production.

Together, we define what the environment or model needs to do, how to measure the result, and when it is ready for deployment.

  1. 1

    Define the task and baseline

    Map data, constraints, and hardware. The current solution becomes the reference for evaluating experiments.

  2. 2

    Build the environment or dataset

    Design RL scenarios and rewards, or curate training and test examples for the SLM.

  3. 3

    Train and evaluate

    Compare versions, analyze failures, and measure quality, cost, and latency against agreed criteria.

  4. 4

    Deploy and monitor

    Integrate the solution, monitor behavior, and plan new cycles based on the results.

Lunar / AI Lab

What stays with your team

  • RL environment or specialized model
  • Datasets, scenarios, and training configuration
  • Baselines and evaluation reports
  • Pipelines, integrations, and documentation
  • Knowledge transfer to your team

Deliverables, responsibilities, and ongoing support are agreed in the project scope.

Your data. Your environment. Your solution.

We design deployment around your requirements: private cloud, on-premise, or isolated environments. Integrations, data access, and model selection are part of the architecture from the start.

  • Private cloud or on-premise
  • Connected to your systems
  • Architecture agreed with your team

Experiment, compare, improve.

Our infrastructure tracks traces, evaluations, and model versions. It supports experiment comparison and failure analysis throughout RL and SLM projects.

Explore the infrastructure
Lunar evaluation interface · example workspace

Before we get started.

Let’s talk

What does an RL environment project deliver?

A training environment with observations, actions, rewards, and scenarios defined for the task. The project can include policy training, evaluation, and integration with your systems, depending on the scope.

What is a Small Language Model in this context?

A compact language model, selected and specialized for tasks in your domain. We define the model, dataset, and fine-tuning or distillation strategy around your quality, latency, cost, and hardware requirements.

Do we need both RL and SLMs?

No. These are two workstreams that can operate independently. We can build an RL environment, specialize an SLM, or combine both approaches when the use case calls for it.

How does a project start?

We discuss the task, available data, and infrastructure. Then we define the scope, baseline, deliverables, and acceptance criteria. Agents and integrations are included when they are part of the application.

What do you need to train?

An environment for learning to decide. A model for mastering a task. Bring your context and let’s define the project.

Let’s discuss your project

RL environments and SLMs, from research to deployment.