All roles

Research · Equity-first · approximately 0.5%

Behavioral Science & Validation Researcher

15–25 hrs / week initially · path to full-timeRemote / flexible

About Aetherya

Aetherya is a frontier cognitive simulation lab building computational models of human perception, behavior and decision-making.

Our research focuses on synthetic humans: persistent, behaviorally grounded agents capable of interpreting information, experiencing uncertainty, forming objections, changing beliefs and making decisions within simulated environments.

We combine foundation models with cognitive architectures, behavioral dynamics, population grounding and multi-agent simulation to study how individuals and groups react before real-world exposure.

Overview

This role owns measurement, experimentation and validation for synthetic behavior. You will design human-versus-synthetic studies, calibration datasets and repeatable benchmarks that test whether Aetherya reproduces meaningful behavioral patterns—not whether its outputs merely sound plausible. Your work will shape model-development priorities, product claims, customer confidence and Aetherya’s external scientific standard.

Your mandate

Create the validation discipline for a new class of behavioral simulation. You will define what credible correspondence between human and synthetic populations means at individual, segment, population and longitudinal levels, then build studies that can detect both improvement and regression.

The objective is not a single headline accuracy score. We need a transparent validity argument: which constructs and decisions are modeled, against which human evidence, under what conditions, with what uncertainty, and where results should not be generalized.

What you will own

  • A multi-level validation framework spanning construct, internal, convergent, discriminant, predictive and external validity
  • Human-versus-synthetic experiments using equivalent stimuli, tasks, exposure conditions and outcome definitions
  • Benchmark suites for preference, recall, trust, persuasion, choice, abandonment, response change and segment differences
  • Sampling plans, power analyses, randomization, controls, exclusions, preregistration and analysis plans
  • Psychometric instruments and behavioral measures with documented reliability and measurement invariance
  • Calibration metrics for distributions, rank ordering, heterogeneity, confidence and behavior over time
  • Bias and failure-mode studies covering agreeableness, rationality, homogeneity, stereotypes, acquiescence and prompt sensitivity
  • Versioned benchmark datasets, reproducible analysis pipelines, scorecards and model-release acceptance criteria

What you will measure

  • Population-level means, proportions, distributions and uncertainty intervals
  • Segment-level direction, magnitude and rank-order agreement
  • Within-person stability and appropriate response to changed evidence or context
  • Calibration, discrimination, sensitivity, specificity and decision-relevant error
  • Variance decomposition across personas, prompts, seeds, models and repeated runs
  • Effect-size recovery rather than significance alone
  • Robustness across wording, modality, ordering, temperature and scenario framing
  • Failure severity based on the downstream decision a customer might make

Example research programs

You might preregister a study in which human participants and synthetic agents evaluate the same product concept, advertisement or risky choice. You would align sampling and exposure, define primary and secondary outcomes, compare full response distributions and segment effects, quantify uncertainty, examine sensitivity to modeling choices and publish the failure cases—not only the strongest agreement.

Another program might reproduce a set of robust effects from behavioral literature, with positive and negative controls, then test whether Aetherya captures effect direction, magnitude, heterogeneity and boundary conditions. A third might construct an adversarial benchmark for agents that are unrealistically agreeable, articulate, rational, stable or demographically stereotyped.

Required research experience

We expect graduate-level capability in experimental design and quantitative inference. Strong candidates will usually be current or completed master’s or PhD researchers, postdoctoral researchers, quantitative user researchers, research scientists, psychometricians or applied statisticians in behavioral science, experimental psychology, cognitive science, behavioral economics, computational social science, HCI or a related discipline.

Conventional job tenure is not required. Relevant experience can come from a thesis, dissertation, laboratory research, registered report, replication, field experiment, research-assistantship, serious quantitative user-research program or independent benchmark project. What matters is evidence that you can design a defensible study, handle messy data, identify threats to inference and communicate results without overstating them.

Minimum bar

  • Independent ownership of at least one substantial quantitative or mixed-method behavioral study
  • Strong command of controls, sampling, randomization, power, effect sizes, missingness and multiple comparisons
  • Ability to distinguish measurement reliability, construct validity and predictive usefulness
  • Experience analyzing participant-level data and reporting uncertainty and practical significance
  • Ability to find confounds, demand characteristics, leakage and alternative explanations before launch
  • Clear writing of protocols, analysis plans, research reports and limitations

Methods & technical profile

  • Advanced working ability in R, Python, Julia, Stata or an equivalent reproducible analysis environment
  • Frequentist and/or Bayesian hierarchical modeling, multilevel analysis and simulation-based power analysis
  • Psychometrics, item-response theory, factor analysis or measurement-invariance testing
  • Distributional comparison, calibration, equivalence testing and reliability analysis
  • Causal inference, experimental design, survey methodology or sequential experimentation
  • Version control, data dictionaries, reproducible pipelines and automated benchmark reporting

Research ethics & governance

  • Design proportionate consent, privacy, retention and de-identification procedures for human studies
  • Identify when institutional ethics review or external research partners are required
  • Document dataset provenance, consent scope, licensing and population limitations
  • Prevent benchmark leakage and separate development sets from held-out evaluation
  • Make negative, null and conflicting findings visible
  • Translate evidence into product claims that remain accurate under scientific and regulatory scrutiny

How you will work

You will work directly with Aetherya’s founder, cognitive simulation researchers and ML/simulation engineers. You will have authority to challenge model claims, stop weak evaluations from becoming marketing evidence and require clearer instrumentation before a mechanism is called validated.

Initial involvement is expected to average approximately 15–25 hours per week. We can accommodate a master’s or PhD program when the candidate can own major studies, meet agreed milestones and address publication, institutional ethics, data ownership and intellectual-property constraints before work begins.

First six months

  • Audit current evidence, metrics and product claims; produce a validation-risk register
  • Define a benchmark taxonomy, reporting standard and minimum evidence levels for major use cases
  • Design and execute one powered human-versus-synthetic pilot with a preregistered analysis plan
  • Build an automated scorecard comparing central tendency, distribution, heterogeneity, calibration and robustness
  • Identify systematic failure modes and convert them into prioritized model-development tickets
  • Prepare a methods report, benchmark release or external paper based on the strongest defensible findings

How success is evaluated

  • Validity, power and reproducibility of studies
  • Ability to detect meaningful failures before customers do
  • Benchmark coverage across constructs, populations, modalities and model versions
  • Clear connection between evaluation findings and model or product decisions
  • Integrity of uncertainty, negative results and limitations reporting
  • Quality of external methods, benchmark and publication outputs

Compensation & path to full-time

This role begins as an equity-only early-team position. Indicative grant: approximately 0.5% of Aetherya, calibrated to contribution, scope and expected long-term involvement, and subject to vesting, cliffs, company approvals and final legal documentation. No cash salary is offered during the initial part-time phase.

The role is designed to transition into a full-time research position once agreed financing or recurring-revenue milestones are reached and both sides confirm long-term fit. At transition, Aetherya intends to offer an above-market full-time cash salary for the relevant location and level. The agreed equity grant remains in place and continues under its vesting terms; it is not exchanged for salary.

Before accepting, candidates will receive written terms covering the initial scope, review milestones, expected transition conditions, vesting, intellectual property, publication, research ethics and termination. Candidates should evaluate the initial equity-only period according to their own financial circumstances.

What to include when applying

  • A CV or academic resume
  • One or two research artifacts: paper, thesis chapter, preregistration, analysis notebook, study protocol or benchmark
  • A short explanation of your personal contribution and the hardest methodological decision in each artifact
  • A brief critique of one way synthetic-behavior systems are commonly evaluated poorly
  • Your realistic weekly availability, program or employment constraints and earliest start date

Interested?

Help define what Aetherya becomes.