Work within authorized data, model goals and compute budgets: pretraining assessment, SFT, preference optimization, RL and private deployment.
Original AI-generated concept scene; not a real customer, device or project result.
Route selection
Compare prompting, RAG, fine-tuning and continued pretraining before training from scratch. Specify parameters, context, tasks and licenses; fitting weights in memory does not establish affordable training.
Data and budget
Track provenance, rights, deduplication, redaction and quality. Separate train, validation and test to prevent leakage. Budget weights, gradients, optimizer, activations, communication and checkpoints; calibrate with small runs.
Training deliverables
Subject to licenses, deliver weights or adapters, configurations, data documentation, logs, evaluation and deployment. Define reproducibility and dependencies; reevaluate model upgrades.
What we can develop
Model and data R&D: architectures, governance, rights and training sets.
Continued pretraining, SFT, LoRA/QLoRA and domain adaptation.
Preference optimization and RL: task-appropriate DPO, reward modeling and GRPO/PPO.
Inference optimization, evaluation, private serving, versioning and safety tests.
Task adaptation and evaluation for text, code, languages and multimodal models.
Annotation, quality filtering, synthetic-data validation and dataset governance.
Distributed training, checkpoints, inference optimization and serving.
Preference data, rewards, safety alignment and controlled-agent training.
From requirements to handover
Requirements and authorization: define the problem, owners, data and interface rights.
Plan and baseline: agree scope, risks, budget, deliverables and acceptance.
Prototype and pilot: validate critical flows and recovery in controlled environments.
Integration and acceptance: review test evidence, not demonstrations alone.
Handover and maintenance: deliver docs, training, access, backups and iteration plans.
Scenario and solution studies
Concept studies illustrate design and acceptance, not completed customer projects.
Domain support-model adaptation
Project context: A base model lacks product terminology and reply format.
Solution approach: Test RAG first, then SFT/LoRA with authorized examples; retain knowledge updates in retrieval.
Acceptance focus: Compare held-out quality, abstention and costs; prevent leakage.
Policy training for verifiable tasks
Project context: A team needs reliable constrained-tool execution.
Solution approach: Define sandbox rewards and safety; compare SFT and GRPO without real funds or production equipment.
Acceptance focus: Check success, violations, reward hacking and cross-task generalization.
Multilingual domain evaluation
Project context: Business quality may vary by language.
Solution approach: Build authorized held-out tasks, compare base/RAG/tuning and report per language.
Acceptance focus: Check coverage, annotation agreement, leakage and failures, not one total score.
Reward research for controlled tool agents
Project context: Tool selection and task completion need verifiable rewards.
Solution approach: Sandbox success, violation and resource rewards; check gaming before small-scale training.
Public case: staged training of a 3B model illustrates data and compute planning; not Xinan work.
Before we start, tell us
Task, base-model license, data sources, samples, GPU budget, serving limits and safety.
Delivery and usage boundaries
Training resources depend on configuration; personal-box inference capability is not training capacity. No claim of general superintelligence or fixed gains.