Laboratory Mozilla Fellowship Research
Paqarina generated tinku
Our model, trained on ninety minutes of Andean archive, produced from noise a sequence that exists in no capture. As far as we could verify, it is the first generative model of an Andean dance.
By Violeta Ayala — Tecnóloga Creativa
Four seconds of dance that nobody danced, produced from noise by a system trained on ninety minutes of Andean archive. As far as we could verify, it is the first generative model of an Andean dance.
The system already knew how to recombine: choose which captured fragment follows which and chain them into something convincing. That is retrieval, not synthesis. Every frame had been danced by a person.
On 5 September the full chain produced a sequence that exists in no capture. A diffusion model sampled a trajectory from noise, another turned it into a body, a third gave it travel across the floor, and the result was mounted on a dancer's skeleton.
It reads. The carriage of the arms, the bend of the elbow, the drop of the step. For once the measurements agree with the eye: lateral elbow separation comes out at 0.277 against the archive's 0.385, and elbow angle at 145 degrees against 128.
What breaks the result is the sliding. The feet move twice as fast as a real dancer's and never slow down. The figure is right and the weight is missing.
What the model does
It learned from 88 motion capture takes made in Cochabamba: tinku, kullawada and breaking. It recognizes, by comparing an unlabelled six second window against the rest of the archive. It recombines, chaining captured fragments into new sequences. And now it generates: it produces four seconds that are in no capture, conditioned on the dance you ask for, with the full body in three dimensions and travel across the floor.
The obvious fix did not work
Sliding has a standard fix in animation. You detect the frames in which a foot is planted and lock it to the floor; the rest of the body moves around it. It is what makes a character walk instead of float.
Detecting it requires a threshold, and published thresholds come from walking data. So we went looking for it in the archive. When the foot is low in a tinku take, how fast does the ankle move?
Horizontal ankle speed, foot low
Frames with the ankle in the lower third of its vertical range. If there were a near still ankle phase separate from the rest, it would appear as a peak beside zero.
Provenance: horizontal ankle speed between consecutive frames at 30 fps, over left and right ankle. The floor is estimated per take as the median of the lowest 5 to 15 per cent of heights, after trimming the extreme 1 per cent, because absolute height varies between sessions. Tinku: 42 takes, 77,592 ankle samples. Breaking from the archive: 6 takes, 18,938 samples. AIST++ breaking: 103 publicly available takes. The vertical axis is normalized to each dance's own maximum, so shapes are comparable and magnitudes are not. Bin widths are those of the original analysis and differ between dances; they are drawn at their true position and width on a common axis.
In tinku the peak is not at zero: it falls between 0.05 and 0.10 metres per second, and the distribution descends continuously from there. In breaking, captured in the same room with the same cameras, the peak sits beside zero and drops sharply. Kullawada, measured on heel and toe and therefore left out of this chart, has the same continuous shape as tinku.
That rules out a limitation of the tracking: the system records a stationary ankle when there is one. And to check it from outside we compared against AIST++, a breaking archive recorded by another team on another continent. Same shape, same peak beside zero.
In these captures we did not find, in tinku, a near still ankle phase separated from the rest of the movement.
It is worth saying what that does not show. A histogram of ankle speed does not prove there is no ground contact: the ankle can displace while the forefoot keeps its support, and the lower third of the range is not the same as touching the floor. What it does show is that the discrete state foot locking needs in order to work does not appear. Locking the foot with inverse kinematics would have imposed a mechanics these data do not support.
It does not walk either
The second assumption to fall was direction. Locomotion models assume a body advances in the direction it faces. We measured where the hip travels relative to the torso.
Direction of travel relative to the torso
Angular distribution over the frames in which the hip displaces. The dashed lines mark the cuts at ±45 and ±135 degrees that define forward, sideways and backward.
Provenance: 37,035 frames from 42 tinku takes in which hip speed exceeds 0.05 m/s. Direction is projected onto the torso's forward and lateral axes, defined by the hip to shoulder vector and the hip to hip vector. The sectors drawn are 30 degrees wide; the percentages are computed on the cuts at ±45 and ±135, which do not coincide with the sector edges.
Where the weight lives
If the ankle never stops, what separates a dancer who looks weighted from a model that does not?
The leg moves against the body's travel. When the hip advances, the low foot displaces in the opposite direction relative to it and the two velocities partly oppose each other. The foot slows against the floor without freezing. It can be measured as the cosine between the two velocities: negative means opposition.
Opposition between hip and leg velocity
Low foot frames in which the cosine between the two velocities is negative.
Provenance: cosine between the horizontal velocity of the hip and that of the ankle relative to the hip, over the 30 per cent of frames with the lowest ankle in each 144 frame window. It measures relative orientation between two velocities, not physical force. The model branch uses the take's real hip trajectory, so it evaluates the generated body's coordination against travel it did not produce.
The distinction changes what needs fixing. Lowering foot speed directly was tried three ways: with a loss function, by moving the model's internal code, and by blending it toward slow examples. None achieved a selective reduction of the low tail. The loss and the blend dragged the high speeds down with it; moving the internal code raised the minimum speed in both directions, which suggests it was pushing the code out of the region where the decoder was trained.
What did not work
Dead ends are most of the work and almost never get published.
The piece that turns the internal code into a body was discarding more than half of what the model knew. Removing some badly chosen loss functions and widening it raised reconstruction from 17 to 45 per cent, measured on takes it had never seen. Four different ways of trying to get past that landed between 45 and 48.
Seven variants over the same internal code
All start from the same frozen representation model and train on the same split.
| Variant | What changes | Reconstruction |
|---|---|---|
| Original | 17 % | |
| Clean and wider | no auxiliary losses | 45 % |
| Stochastic | objective that admits variety | 45 % |
| Temporal | sees the whole window | 47 % |
| Slow foot | loss on the low tail | 47 % |
| With direction | told where the body travels | 48 % |
| With direction, variant | another coordinate system | 47 % |
Provenance: the reconstruction column is the variance of a movement feature vector, the three dimensional range covered by each joint within the window, expressed as a fraction of the archive's. It is not R². Measured on 18 held out takes; no window from a take appears in both training and evaluation.
And four times the instrument was wrong
In a single day, four measurement errors each produced a false conclusion before being found. Comparing 144 generated frames against 1700 real ones, and concluding the generated arm was frozen. Sampling windows in file order, so a dataset's figures swung by a factor of three depending on how many were taken. Rotating a pose into a coordinate system it was already in. Comparing a branch that had been through post processing against one that had not.
The most expensive one came earlier: two days chasing an impossible elbow inside the model, until a real take, run through the same Blender skeleton, bent exactly the same way. The defect was in the rigging.
The archive is the instrument. It needs checking as often as the model does.
Whose archive this is
Published dance generators train on tens of hours of street or commercial material, on datasets assembled in studios. This one trained on ninety minutes. That difference is technical and can be argued about.
The other one cannot. The archive was made by the people who dance, in their own city, and it belongs to them. The labels, the cleaning of the tracking and the decision about which takes go in and which stay out passed through hands that know the dance. These measurements start from that knowledge and serve to interrogate what the model preserves, not to validate it.
What comes next is a concrete test: whether the opposition between hip and leg can be recovered from the representation we already have, or whether the model has to be trained to preserve weight transfer from the start. And we keep recording. In the tests run so far, bodies move the needle more than changes of architecture do.
Paqarina is a motion capture archive of Andean dance and a computational creativity lab in Cochabamba. The figures in this piece come from 88 takes captured between 2025 and 2026. The breaking comparison material comes from AIST++, which is publicly available. The claim of being first was checked by literature search in September 2026 and is limited to Andean dance: there is published generative work on Bharatanatyam, the classical dance of southern India, producing key postures and short sequences from an archive of ten dancers.
The dance belongs to those who dance it.
