# Pre-training Transformer model on encrypted data

**URL:** https://openfhe.discourse.group/t/pre-training-transformer-model-on-encrypted-data/1276
**Category:** FHE Use Cases
**Created:** [May 17, 2024, 12:26pm UTC](https://openfhe.discourse.group/t/pre-training-transformer-model-on-encrypted-data/1276 "2024-05-17T12:26:36Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![Salay](https://avatars.discourse-cdn.com/v4/letter/s/a698b9/32.png) [@Salay](https://openfhe.discourse.group/u/Salay)
#### Post date: [May 17, 2024, 12:26pm UTC](https://openfhe.discourse.group/t/pre-training-transformer-model-on-encrypted-data/1276/1 "2024-05-17T12:26:36Z")

</div>

With classical Transformer, we transform each word in sentence to an embedding vector d dimesional. So we get a matrix X of nxd dimensional for a sentence with n words. But each vector will be transformed on polynomial form in CKKS scheme. And we get a sequence of polynomial for this sentence. How do I calculate the attentions scores for this polynomial sequence?

---

<div class="post-metadata">

### Author: ![Caesar](https://yyz1.discourse-cdn.com/flex031/user_avatar/openfhe.discourse.group/caesar/32/63_2.png) [@Caesar](https://openfhe.discourse.group/u/Caesar)
#### Post date: [May 17, 2024, 7:25pm UTC](https://openfhe.discourse.group/t/pre-training-transformer-model-on-encrypted-data/1276/2 "2024-05-17T19:25:09Z")

</div>

You should not view the encoded values as polynomials, but rather as vectors.

Assuming that d \< `num_slots`, you can think of X as n of d-dimensional vectors. When you use CKKS operations such as adding or multiplying two encoded vectors (whether these vectors are plaintexts or ciphertexts), the result is an encoded vector of point-wise addition or multiplication.

In the attention layer, you need to compute the matrix product of X with the query, key, and value weight matrices. For the attention score, you need to process the query and key results and sqrt(d\_K). Lastly, you need to compute the `softmax_max` function which can be done by polynomial approximation. This all can be done in CKKS.

You can have a look at this [paper](https://eprint.iacr.org/2024/136.pdf) which evaluated a simple transformer in CKKS.

---

<div class="post-metadata">

### Author: ![Salay](https://avatars.discourse-cdn.com/v4/letter/s/a698b9/32.png) [@Salay](https://openfhe.discourse.group/u/Salay)
#### Post date: [May 18, 2024, 2:51pm UTC](https://openfhe.discourse.group/t/pre-training-transformer-model-on-encrypted-data/1276/3 "2024-05-18T14:51:33Z")

</div>

It means I have wrong. I have encoded each row of the embedding Vector to a polynome
