# CKKS Memory Issue for Matrix Multiplication by 256x256

**URL:** <https://openfhe.discourse.group/t/ckks-memory-issue-for-matrix-multiplication-by-256x256/1928>\
**Category:** Library Questions\
**Created:** [March 14, 2025, 4:18pm UTC](https://openfhe.discourse.group/t/ckks-memory-issue-for-matrix-multiplication-by-256x256/1928 "2025-03-14T16:18:43Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![seyda](https://yyz1.discourse-cdn.com/flex031/user_avatar/openfhe.discourse.group/seyda/32/397_2.png) [@seyda](https://openfhe.discourse.group/u/seyda)\
**Post date:** [March 14, 2025, 4:18pm UTC](https://openfhe.discourse.group/t/ckks-memory-issue-for-matrix-multiplication-by-256x256/1928/1 "2025-03-14T16:18:43Z")

</div>

Hi, I am trying to understand the implementation of the Matrix Multiplication at the [polyakov-matrix-mult branch](https://github.com/openfheorg/openfhe-development/blob/polyakov-matrix-mult/src/pke/examples/matrix-mult-examples.cpp) upon a previous discussion at the [Matrix Multiplication with Diagonal Packing - #2 by ypolyakov](https://openfhe.discourse.group/t/matrix-multiplication-with-diagonal-packing/1893/2).

Looking at the rotations in the [square matrix multiplication algorithm](https://github.com/openfheorg/openfhe-development/blob/d681986092dee778e34764aacee082730ca7fdd5/src/pke/examples/matrix-mult-examples.cpp#L604), Is my understanding correct: For a 128x128 matrix and a parameter set (log N = 16, loq PQ = 1728, L = 20, dnum = 3), for each distinct rotation index we need a key of size ~ 83 MB (dnum \* 2 \* N \* log PQ / 8). We need rotations by i and -i, i, and i\*rowsize (1, … , 127), where i = (1,…,127). This adds up to 127 \* 3 rotation keys, each with size ~ 83 MB, adding up to 31.6 GB.

My question: Is there any way to increase the log N = 17 and perform 256x256 matrix multiplications without memory issues on a GPU?

---

<div class="post-metadata">

**Author:** ![Caesar](https://yyz1.discourse-cdn.com/flex031/user_avatar/openfhe.discourse.group/caesar/32/63_2.png) [@Caesar](https://openfhe.discourse.group/u/Caesar)\
**Post date:** [March 14, 2025, 6:46pm UTC](https://openfhe.discourse.group/t/ckks-memory-issue-for-matrix-multiplication-by-256x256/1928/2 "2025-03-14T18:46:53Z")

</div>

It seems you have your own GPU FHE library. Do you need to keep all the keys in the GPU memory? Maybe you can push them to GPU memory when needed or prefetch them slightly a head of the time they are needed.

Looking at this from a different angle, would it make sense to decompose your 256x256 big matrix into 4 128x128 sub-matrices, and compute the big matrix product using 8 sub-matrix products and 4 sub-matrix additions?

---

<div class="post-metadata">

**Author:** ![seyda](https://yyz1.discourse-cdn.com/flex031/user_avatar/openfhe.discourse.group/seyda/32/397_2.png) [@seyda](https://openfhe.discourse.group/u/seyda)\
**Post date:** [March 14, 2025, 6:54pm UTC](https://openfhe.discourse.group/t/ckks-memory-issue-for-matrix-multiplication-by-256x256/1928/3 "2025-03-14T18:54:34Z")

</div>

Hi, I wasn’t running any code on the GPU at the moment, but planning for ahead on our library.

Right now I am following your suggestion, decomposing the matrix. I will consider the prefetching part, thank you so much!

Do you think my calculations are correct (especially in terms of the required number of rotation keys) and reflect the correct metadata size, or did I overdo some part and incorrectly got a high value?

---

<div class="post-metadata">

**Author:** ![Caesar](https://yyz1.discourse-cdn.com/flex031/user_avatar/openfhe.discourse.group/caesar/32/63_2.png) [@Caesar](https://openfhe.discourse.group/u/Caesar)\
**Post date:** [March 14, 2025, 8:14pm UTC](https://openfhe.discourse.group/t/ckks-memory-issue-for-matrix-multiplication-by-256x256/1928/4 "2025-03-14T20:14:13Z")

</div>

Your estimations look good to me.
