lunar
Let’s talk
Lunar · Enterprise AI Lab

RL environments.
Small language models.

We build reinforcement learning environments and train compact language models for your business tasks. From experiment design to deployment, with your data and your team.

RL environmentsSmall Language ModelsDeployment in your infrastructure
RL environments & SLMs

Your training environment. Your specialized model.

RL environments and Small Language Model (SLM) projects, delivered independently or combined to fit the challenge. Each workstream has defined deliverables and evaluation criteria.

RL environments

The right context for learning to decide.

We turn your business rules, data, and interactions into reinforcement learning environments. We define observations, actions, rewards, and scenarios to train and evaluate agents or policies.

  • Versioned environments and training scenarios
  • Task-specific rewards and constraints
  • Policy evaluation against a baseline
Lunar model evolution workspace
Environment design

Your business states, actions, and scenarios.

We model available observations, permitted actions, and state transitions. Training and test scenarios reflect operational rules and data, including failure cases and constraints.

Reward engineering

A reward that reflects the task.

We define reward signals and penalties to guide learning. We evaluate unwanted behavior and compare policies with existing rules before deployment.

Lunar automatic evaluations
Small Language Models

Compact models. Well-defined tasks.

We select and specialize smaller language models for your domain. Through curation, fine-tuning, and distillation, we target the right balance of quality, cost, and latency for your tasks.

  • Curated datasets and separate test sets
  • Model fine-tuning and distillation
  • Evaluation and serving on the selected hardware
Fine-tuning and distillation

Specialize the model for your domain.

We curate examples and reference-model outputs to train the SLM. Versions are compared on a separate test set and improvements are validated under the project’s operating conditions.

Evaluation and serving

Quality measured in your environment.

The model needs to meet the task requirements on the available hardware. We evaluate results before integrating the SLM into your systems.

  • Task-level quality and failure analysis
  • Cost, latency, and memory usage
  • Versioning and production monitoring
Owned AI Infrastructure

Your data. Your models. Your servers.

A full AI stack deployed inside your perimeter — gateway, traces, training, serving. Compliance-ready. Open-source. No data leaves your environment.

Lunar model policies
Full-stack AI inside your perimeter

Your data never leaves your environment.

A complete AI stack — gateway, tracing, training, serving — deployed inside your own infrastructure. On-premise, private cloud, or air-gapped. No outbound data, ever.

  • On-premise, private cloud, or air-gapped
  • LGPD, GDPR, HIPAA, SOC 2 ready
  • Open-source — no vendor lock-in, ever

On-premise or private cloud

Deploy to AWS, GCP, Azure, or bare metal. You choose where the stack lives.

Air-gapped available

For regulated environments with strict network isolation. No outbound traffic.

Open-source, MIT license

Full codebase visibility. Fork it, audit it, extend it. No lock-in.

LGPD · GDPR · HIPAA · SOC 2

Pre-built compliance postures for each framework. Your team stays in control.

Industries

Where RL environments and SLMs fit.

Example tasks to start the conversation. The initial assessment identifies an approach that fits the project’s data and constraints.

Operations & Logistics

Routing, scheduling, predictive maintenance.

Finance & Risk

Pricing, credit scoring, fraud detection.

Healthcare

Triage, decision support, clinical summaries.

Legal

Contracts, compliance, knowledge agents.

Engineering

Coding agents, PR review, build intelligence.

Customer Experience

Support, FAQ, lead qualification.

Data & Analytics

SQL, extraction, executive insights.

Regulated & Government

Air-gapped, redacted, compliance-first.

Services

From dataset to deployment, with your team.

Together, we define what the environment or model needs to do, how to measure the result, and when it is ready for deployment.

Training data

Curation of examples, trajectories, and quality criteria for each experiment.

Reproducible evaluations

Baselines, test scenarios, and reports to decide what goes into production.

Infrastructure and serving

Training and inference pipelines sized for your environment.

Operational integration

Models and policies connected to agents, tools, and internal systems.

What do you need to train?

An environment for learning to decide. A model for mastering a task. Bring your context and let’s define the project.

RL environments and SLMs, from research to deployment.