Work within authorized data, model goals and compute budgets: pretraining assessment, SFT, preference optimization, RL and private deployment.
Original AI-generated concept scene; not a real customer, device or project result.
Data and collection definitions
Inventory licenses, data, annotation, preferences, rewards and held-out tests. Count samples/tokens, lengths, languages, duplicates and privacy; check overlap. Prioritize rights and quality over scale and verify synthetic bias/contamination.
Interfaces and integration
Training covers GPU topology, communication, storage, checkpoints and tracking. Assess serving batch, cache, context and concurrency separately. Bind model, adapters, data and environment versions; do not reuse unauthorized customer data.
Operations and handover
Report actual spend, logs, held-out evaluation, failures, safety and licenses. Exercise resume and rollback; monitor drift, latency and cost. Successful training does not guarantee business success; define launch scope and owners.
What we can develop
Model and data R&D: architectures, governance, rights and training sets.
Continued pretraining, SFT, LoRA/QLoRA and domain adaptation.
Preference optimization and RL: task-appropriate DPO, reward modeling and GRPO/PPO.
Inference optimization, evaluation, private serving, versioning and safety tests.
Task adaptation and evaluation for text, code, languages and multimodal models.
Annotation, quality filtering, synthetic-data validation and dataset governance.
Distributed training, checkpoints, inference optimization and serving.
Preference data, rewards, safety alignment and controlled-agent training.
From requirements to handover
Requirements and authorization: define the problem, owners, data and interface rights.
Plan and baseline: agree scope, risks, budget, deliverables and acceptance.
Prototype and pilot: validate critical flows and recovery in controlled environments.
Integration and acceptance: review test evidence, not demonstrations alone.
Handover and maintenance: deliver docs, training, access, backups and iteration plans.
Scenario and solution studies
Concept studies illustrate design and acceptance, not completed customer projects.
Domain support-model adaptation
Project context: A base model lacks product terminology and reply format.
Solution approach: Test RAG first, then SFT/LoRA with authorized examples; retain knowledge updates in retrieval.
Acceptance focus: Compare held-out quality, abstention and costs; prevent leakage.
Policy training for verifiable tasks
Project context: A team needs reliable constrained-tool execution.
Solution approach: Define sandbox rewards and safety; compare SFT and GRPO without real funds or production equipment.
Acceptance focus: Check success, violations, reward hacking and cross-task generalization.
Multilingual domain evaluation
Project context: Business quality may vary by language.
Solution approach: Build authorized held-out tasks, compare base/RAG/tuning and report per language.
Acceptance focus: Check coverage, annotation agreement, leakage and failures, not one total score.
Reward research for controlled tool agents
Project context: Tool selection and task completion need verifiable rewards.
Solution approach: Sandbox success, violation and resource rewards; check gaming before small-scale training.
Public case: staged training of a 3B model illustrates data and compute planning; not Xinan work.
Before we start, tell us
Task, base-model license, data sources, samples, GPU budget, serving limits and safety.
Delivery and usage boundaries
Training resources depend on configuration; personal-box inference capability is not training capacity. No claim of general superintelligence or fixed gains.