{"id":443737,"date":"2025-01-01T09:00:19","date_gmt":"2025-01-01T09:00:19","guid":{"rendered":"http:\/\/savepearlharbor.com\/?p=443737"},"modified":"-0001-11-30T00:00:00","modified_gmt":"-0001-11-29T21:00:00","slug":"","status":"publish","type":"post","link":"https:\/\/savepearlharbor.com\/?p=443737","title":{"rendered":"<span>\u041f\u0438\u0448\u0435\u043c \u0441\u0432\u043e\u0439 PyTorch \u043d\u0430 NumPy. \u0424\u0438\u043d\u0430\u043b. \u0417\u0430\u043f\u0443\u0441\u043a\u0430\u0435\u043c GPT-2<\/span>"},"content":{"rendered":"<div><!--[--><!--]--><\/div>\n<div id=\"post-content-body\">\n<div>\n<div class=\"article-formatted-body article-formatted-body article-formatted-body_version-2\">\n<div xmlns=\"http:\/\/www.w3.org\/1999\/xhtml\">\n<p><strong>PyTorch<\/strong>\u00a0\u2014 \u044d\u0442\u043e \u043c\u043e\u0449\u043d\u044b\u0439 \u0438 \u0433\u0438\u0431\u043a\u0438\u0439 \u0444\u0440\u0435\u0439\u043c\u0432\u043e\u0440\u043a \u0434\u043b\u044f \u043c\u0430\u0448\u0438\u043d\u043d\u043e\u0433\u043e \u043e\u0431\u0443\u0447\u0435\u043d\u0438\u044f, \u0448\u0438\u0440\u043e\u043a\u043e \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u0435\u043c\u044b\u0439 \u0434\u043b\u044f \u0441\u043e\u0437\u0434\u0430\u043d\u0438\u044f \u043d\u0435\u0439\u0440\u043e\u043d\u043d\u044b\u0445 \u0441\u0435\u0442\u0435\u0439. \u041e\u043d \u043e\u0441\u043e\u0431\u0435\u043d\u043d\u043e \u043f\u043e\u043f\u0443\u043b\u044f\u0440\u0435\u043d \u0431\u043b\u0430\u0433\u043e\u0434\u0430\u0440\u044f \u043f\u0440\u043e\u0441\u0442\u043e\u0442\u0435 \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u043d\u0438\u044f, \u0434\u0438\u043d\u0430\u043c\u0438\u0447\u0435\u0441\u043a\u0438\u043c \u0432\u044b\u0447\u0438\u0441\u043b\u0438\u0442\u0435\u043b\u044c\u043d\u044b\u043c \u0433\u0440\u0430\u0444\u0430\u043c \u0438 \u0431\u043e\u0433\u0430\u0442\u043e\u0439 \u044d\u043a\u043e\u0441\u0438\u0441\u0442\u0435\u043c\u0435 \u0438\u043d\u0441\u0442\u0440\u0443\u043c\u0435\u043d\u0442\u043e\u0432 \u0434\u043b\u044f \u043e\u0431\u0443\u0447\u0435\u043d\u0438\u044f \u043c\u043e\u0434\u0435\u043b\u0435\u0439. \u0414\u043b\u044f \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u043d\u0438\u044f \u044d\u0442\u043e\u0433\u043e \u0444\u0440\u0435\u0439\u043c\u0432\u043e\u0440\u043a\u0430, \u0447\u0430\u0441\u0442\u043e \u0434\u043e\u0441\u0442\u0430\u0442\u043e\u0447\u043d\u043e \u043f\u043e\u0432\u0435\u0440\u0445\u043d\u043e\u0441\u0442\u043d\u043e \u043f\u043e\u043d\u0438\u043c\u0430\u0442\u044c \u0440\u0430\u0431\u043e\u0442\u0443 \u0430\u043b\u0433\u043e\u0440\u0438\u0442\u043c\u043e\u0432 \u043c\u0430\u0448\u0438\u043d\u043d\u043e\u0433\u043e \u043e\u0431\u0443\u0447\u0435\u043d\u0438\u044f.<\/p>\n<p>\u041d\u043e \u0410\u043d\u0434\u0440\u0435\u0439 \u041a\u0430\u0440\u043f\u0430\u0442\u044b, \u0438\u0437\u0432\u0435\u0441\u0442\u043d\u044b\u0439 \u0438\u0441\u0441\u043b\u0435\u0434\u043e\u0432\u0430\u0442\u0435\u043b\u044c \u0432 \u043e\u0431\u043b\u0430\u0441\u0442\u0438 \u0418\u0418, \u0441\u0447\u0438\u0442\u0430\u0435\u0442, \u0447\u0442\u043e \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044f \u0430\u043b\u0433\u043e\u0440\u0438\u0442\u043c\u043e\u0432 \u0441 \u043d\u0443\u043b\u044f \u043f\u043e\u0437\u0432\u043e\u043b\u044f\u0435\u0442 \u043f\u043e\u043d\u044f\u0442\u044c \u0438\u0445 \u0441\u0443\u0442\u044c \u0438 \u0434\u0435\u0442\u0430\u043b\u0438 \u0440\u0430\u0431\u043e\u0442\u044b, \u0447\u0442\u043e \u0441\u043b\u043e\u0436\u043d\u043e \u043e\u0441\u043e\u0437\u043d\u0430\u0442\u044c, \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044f \u0442\u043e\u043b\u044c\u043a\u043e \u0433\u043e\u0442\u043e\u0432\u044b\u0435 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0438. \u042d\u0442\u043e \u043f\u043e\u043c\u043e\u0433\u0430\u0435\u0442 \u0440\u0430\u0437\u0432\u0438\u0442\u044c \u0438\u043d\u0442\u0443\u0438\u0446\u0438\u044e \u0434\u043b\u044f \u0434\u0430\u043b\u044c\u043d\u0435\u0439\u0448\u0435\u0433\u043e \u043f\u0440\u0438\u043c\u0435\u043d\u0435\u043d\u0438\u044f \u0438 \u0443\u043b\u0443\u0447\u0448\u0435\u043d\u0438\u044f \u043c\u0435\u0442\u043e\u0434\u043e\u0432. \u0410\u043d\u0434\u0440\u0435\u0439 \u043f\u043e\u0441\u0432\u044f\u0449\u0430\u0435\u0442 \u043c\u043d\u043e\u0433\u043e \u0441\u043e\u0431\u0441\u0442\u0432\u0435\u043d\u043d\u043e\u0433\u043e \u0432\u0440\u0435\u043c\u0435\u043d\u0438, \u0447\u0442\u043e\u0431\u044b \u043e\u0431\u044a\u044f\u0441\u043d\u044f\u0442\u044c \u043a\u043b\u044e\u0447\u0435\u0432\u044b\u0435 \u043f\u0440\u0438\u043d\u0446\u0438\u043f\u044b \u0440\u0430\u0431\u043e\u0442\u044b \u043d\u0435\u0439\u0440\u043e\u0441\u0435\u0442\u0435\u0439 \u0432 \u0441\u0432\u043e\u0438\u0445\u00a0<a href=\"https:\/\/karpathy.ai\/\" rel=\"noopener noreferrer nofollow\">\u0431\u043b\u043e\u0433\u0430\u0445<\/a>\u00a0\u0438 \u043d\u0430 \u0441\u0432\u043e\u0451\u043c\u00a0<a href=\"https:\/\/www.youtube.com\/@AndrejKarpathy\" rel=\"noopener noreferrer nofollow\">\u044e\u0442\u0443\u0431-\u043a\u0430\u043d\u0430\u043b\u0435<\/a>. \u041e\u043d \u0442\u0430\u043a\u0436\u0435 \u043d\u0435 \u0440\u0430\u0437 \u043f\u043e\u0434\u0447\u0435\u0440\u043a\u0438\u0432\u0430\u043b, \u0447\u0442\u043e \u043d\u0430 \u0435\u0433\u043e \u043a\u0443\u0440\u0441\u0435 \u0432\u00a0<a href=\"https:\/\/cs.stanford.edu\/people\/karpathy\/\" rel=\"noopener noreferrer nofollow\">C\u0442\u044d\u043d\u0444\u043e\u0440\u0434\u0435<\/a>\u00a0\u0435\u0441\u0442\u044c \u0437\u0430\u0434\u0430\u0447\u0438 \u043f\u043e \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u0438 \u0440\u0430\u0437\u043b\u0438\u0447\u043d\u044b\u0445 \u0430\u043b\u0433\u043e\u0440\u0438\u0442\u043c\u043e\u0432, \u043d\u0430\u043f\u0440\u0438\u043c\u0435\u0440, \u043e\u0431\u0440\u0430\u0442\u043d\u043e\u0435 \u0440\u0430\u0441\u043f\u0440\u043e\u0441\u0442\u0440\u0430\u043d\u0435\u043d\u0438\u0435.<\/p>\n<p>\u042f \u0445\u043e\u0442\u0435\u043b \u0431\u044b \u043f\u043e\u0441\u0432\u044f\u0442\u0438\u0442\u044c \u0434\u0430\u043d\u043d\u0443\u044e \u0441\u0442\u0430\u0442\u044c\u044e \u044d\u0442\u043e\u0439 \u0438\u0434\u0435\u0438, \u043f\u043e\u0442\u043e\u043c\u0443 \u0447\u0442\u043e \u043c\u043d\u0435 \u0441\u0430\u043c\u043e\u043c\u0443 \u043e\u0441\u043e\u0431\u0435\u043d\u043d\u043e \u0438\u043d\u0442\u0435\u0440\u0435\u0441\u043d\u043e \u043a\u043e\u043f\u0430\u0442\u044c\u0441\u044f \u0432 \u0430\u043b\u0433\u043e\u0440\u0438\u0442\u043c\u0430\u0445 \u0433\u043b\u0443\u0431\u043e\u043a\u043e\u0433\u043e \u043e\u0431\u0443\u0447\u0435\u043d\u0438\u044f. \u042d\u0442\u0430 \u0441\u0442\u0430\u0442\u044c\u044f \u043f\u0440\u043e\u0434\u043e\u043b\u0436\u0435\u043d\u0438\u0435\u00a0<a href=\"https:\/\/habr.com\/ru\/articles\/870426\/\" rel=\"noopener noreferrer nofollow\">\u0442\u0440\u0435\u0442\u044c\u0435\u0439 \u0441\u0442\u0430\u0442\u044c\u0438<\/a><\/p>\n<p>\u0418\u0442\u0430\u043a, \u043a\u0430\u043a \u044f \u0438 \u0441\u043a\u0430\u0437\u0430\u043b \u0432 \u043f\u0440\u0435\u0434\u044b\u0434\u0443\u0449\u0435\u0439 \u0441\u0442\u0430\u0442\u044c\u0435, \u0443 \u043d\u0430\u0441 \u0434\u043e\u0441\u0442\u0430\u0442\u043e\u0447\u043d\u043e \u0437\u043d\u0430\u043d\u0438\u0439, \u0447\u0442\u043e\u0431\u044b \u0441\u043e\u0431\u0440\u0430\u0442\u044c \u0438\u0445 \u0432 \u0446\u0435\u043b\u0443\u044e \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0443. \u0422\u0430\u043a \u044f \u0438 \u0441\u0434\u0435\u043b\u0430\u043b, \u0440\u0435\u0430\u043b\u0438\u0437\u043e\u0432\u0430\u0432 <a href=\"https:\/\/github.com\/TimaGitHub\/pycandle-2\" rel=\"noopener noreferrer nofollow\">pycandle<\/a>.<br \/>\u0422\u0430\u043c \u043d\u0430\u0445\u043e\u0434\u0438\u0442\u044c\u0441\u044f \u0432\u0441\u0451 \u0442\u043e, \u0447\u0442\u043e \u043c\u044b \u0438\u0437\u0443\u0447\u0430\u043b\u0438 \u043d\u0430 \u043f\u0440\u043e\u0442\u044f\u0436\u0435\u043d\u0438\u0438 3 \u0447\u0430\u0441\u0442\u0435\u0439. \u0412 \u044d\u0442\u043e\u0439 \u0447\u0430\u0441\u0442\u0438 \u043c\u044b \u0431\u0443\u0434\u0435\u043c \u043f\u0438\u0441\u0430\u0442\u044c \u0438\u043d\u0444\u0435\u0440\u0435\u043d\u0441 \u043a\u043e\u0434 \u0434\u043b\u044f <strong>GPT2<\/strong> \u043d\u0430 \u0441\u043e\u0431\u0441\u0442\u0432\u0435\u043d\u043d\u043e\u0439 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0435!<br \/>\u041d\u0430\u0447\u043d\u0451\u043c \u0441 \u0438\u043c\u043f\u043e\u0440\u0442\u0430 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a<\/p>\n<pre><code class=\"python\">import torch import torch.nn as nn<\/code><\/pre>\n<p>\u041e\u0439, \u043d\u0435 \u0442\u043e. \u042f \u0438\u043c\u0435\u043b \u0432 \u0432\u0438\u0434\u0443<\/p>\n<pre><code class=\"python\">import candle import candle.nn as nn<\/code><\/pre>\n<p>\u0422\u0430\u043a\u0436\u0435 \u0432\u043e\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u0435\u043c\u0441\u044f \u0442\u043e\u043a\u0435\u043d\u0438\u0437\u0430\u0442\u043e\u0440\u043e\u043c \u043e\u0442 OpenAI, \u043a\u043e\u0442\u043e\u0440\u044b\u0439 \u043e\u043d\u0438 \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044e\u0442 \u0432 \u0441\u0432\u043e\u0438\u0445 \u043c\u043e\u0434\u0435\u043b\u044f\u0445. \u0421\u0434\u0435\u043b\u0430\u0435\u043c \u0432\u0441\u0435 \u043d\u0435\u043e\u0431\u0445\u043e\u0434\u0438\u043c\u044b\u0435 \u0438\u043c\u043f\u043e\u0440\u0442\u044b<\/p>\n<pre><code class=\"python\">import tiktoken from candle import Tensor from dataclasses import dataclass<\/code><\/pre>\n<p>\u041e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u043c \u043a\u043e\u043d\u0444\u0438\u0433\u0443\u0440\u0430\u0446\u0438\u044e \u043d\u0430\u0448\u0435\u0439 \u043c\u043e\u0434\u0435\u043b\u0438 \u0441 \u043f\u043e\u043c\u043e\u0449\u044c\u044e <code>dataclasses<\/code><\/p>\n<pre><code class=\"python\">@dataclass class GPTConfig:     block_size: int = 1024 # \u0440\u0430\u0437\u043c\u0435\u0440 \u043e\u043a\u043d\u0430 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u04d9     vocab_size: int = 50257 # \u0440\u0430\u0437\u043c\u0435\u0440 \u0441\u043b\u043e\u0432\u0430\u0440\u044f BPE     n_layer: int = 12 # \u043a\u043e\u043b\u0438\u0447\u0435\u0441\u0442\u0432\u043e \u0441\u043b\u043e\u0451\u0432     n_head: int = 12 # \u043a\u043e\u043b\u0438\u0447\u0435\u0441\u0442\u0432\u043e \u0433\u043e\u043b\u043e\u0432 \u0432 \u043c\u0435\u0445\u0430\u043d\u0438\u0437\u043c\u0435 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u044f     n_embd: int = 768 # \u0440\u0430\u0437\u043c\u0435\u0440\u043d\u043e\u0441\u0442\u044c \u0432\u0435\u043a\u0442\u043e\u0440\u0430 \u044d\u043c\u0431\u0435\u0434\u0434\u0438\u043d\u0433\u043e\u0432<\/code><\/pre>\n<p>\u041e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u043c \u043b\u0438\u043d\u0435\u0439\u043d\u044b\u0439 \u0441\u043b\u043e\u0439 \u043c\u043e\u0434\u0435\u043b\u0438<\/p>\n<pre><code class=\"python\">class MLP(nn.Module):     def __init__(self, config):         super().__init__()         self.c_fc = nn.Linear(config.n_embd, 4 * config.n_embd)         self.gelu = nn.GeLU()         self.c_proj = nn.Linear(4 * config.n_embd, config.n_embd)      def forward(self, x):         x = self.c_fc(x)         x = self.gelu(x)         return self.c_proj(x)<\/code><\/pre>\n<p><strong>\u041e\u0431\u0440\u0430\u0442\u0438\u0442\u0435 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u0435<\/strong> \u0442\u0443\u0442 <code>candle.nn<\/code>, \u0430 <strong>\u043d\u0435<\/strong> <code>torch.nn<\/code>. \u0421 \u043b\u0438\u043d\u0435\u0439\u043d\u044b\u043c\u0438 \u0441\u043b\u043e\u044f\u043c\u0438 \u043c\u044b \u0443\u0436\u0435 \u0437\u043d\u0430\u043a\u043e\u043c\u044b, \u0430 <code>GeLU()<\/code>, \u044d\u0442\u043e \u0447\u0442\u043e \u0442\u0430\u043a\u043e\u0435? <\/p>\n<figure class=\"full-width\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/habrastorage.org\/r\/w1560\/getpro\/habr\/upload_files\/4c5\/7e8\/6ae\/4c57e86aeb4ea92fd543a8942c5e92c2.png\" width=\"850\" height=\"533\" data-src=\"https:\/\/habrastorage.org\/getpro\/habr\/upload_files\/4c5\/7e8\/6ae\/4c57e86aeb4ea92fd543a8942c5e92c2.png\"\/><\/figure>\n<p>\u0418\u0441\u0441\u043b\u0435\u0434\u043e\u0432\u0430\u0442\u0435\u043b\u0438 \u0438\u0437 <code>OpenAI<\/code>, \u0432\u044b\u044f\u0441\u043d\u0438\u043b\u0438, \u0447\u0442\u043e \u0432 \u0438\u0445 \u043c\u043e\u0434\u0435\u043b\u044f\u0445 \u0442\u0430\u043a\u0430\u044f \u0444\u0443\u043d\u043a\u0446\u0438\u044f \u0440\u0430\u0431\u043e\u0442\u0430\u0435\u0442 \u043b\u0443\u0447\u0448\u0435 \u0432\u0441\u0435\u0433\u043e, \u043f\u043e\u0441\u043c\u043e\u0442\u0440\u0438\u043c \u043d\u0430 \u0435\u0451 \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044e<\/p>\n<figure class=\"full-width\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/habrastorage.org\/r\/w1560\/getpro\/habr\/upload_files\/226\/b4c\/ff1\/226b4cff1dbcbac0b184d2e09f796fd6.png\" width=\"922\" height=\"354\" data-src=\"https:\/\/habrastorage.org\/getpro\/habr\/upload_files\/226\/b4c\/ff1\/226b4cff1dbcbac0b184d2e09f796fd6.png\"\/><\/figure>\n<pre><code class=\"python\">class GeLU:     def __init__(self):         Parameter([self,[]])      def __call__(self, x):         return Tensor.gelu(x)         #return 0.5 * x * (1 + Tensor.tanh(0.79788456 * (x + 0.044715 * (x ** 3))))<\/code><\/pre>\n<p>\u042f \u0437\u0430\u043a\u043e\u043c\u043c\u0435\u043d\u0442\u0438\u0440\u043e\u0432\u0430\u043b \u0432\u0442\u043e\u0440\u0443\u044e \u0441\u0442\u0440\u043e\u0447\u043a\u0443 \u0438 \u0432\u043c\u0435\u0441\u0442\u043e \u044d\u0442\u043e\u0433\u043e \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044e \u0432\u0441\u0442\u0440\u043e\u0435\u043d\u043d\u044b\u0439 \u043c\u0435\u0442\u043e\u0434 <code>Tensor.gelu()<\/code>, \u043d\u043e \u043d\u0438\u043a\u0442\u043e \u043d\u0435 \u0437\u0430\u043f\u0440\u0435\u0449\u0430\u0435\u0442 \u0438 \u0447\u0435\u0441\u0442\u043d\u043e \u0441\u0447\u0438\u0442\u0430\u0442\u044c \u043f\u043e \u0444\u043e\u0440\u043c\u0443\u043b\u0435. \u0417\u0430\u0433\u043b\u044f\u043d\u0435\u043c \u0432 \u044d\u0442\u043e\u0442 \u043c\u0435\u0442\u043e\u0434<\/p>\n<pre><code class=\"python\">from scipy.stats import norm @classmethod def gelu(cls, x): return x * norm.cdf(x.value) # \u0444\u043e\u0440\u043c\u0443\u043b\u0430 \u0438\u0437 \u043a\u0430\u0440\u0442\u0438\u043d\u043a\u0438<\/code><\/pre>\n<p>\u0422\u0435\u043f\u0435\u0440\u044c \u0432 \u043b\u0438\u043d\u0435\u0439\u043d\u043e\u043c \u0441\u043b\u043e\u0435 \u043d\u0430\u043c \u0432\u0441\u0435 \u043f\u043e\u043d\u044f\u0442\u043d\u043e, \u0438\u0434\u0451\u043c \u0434\u0430\u043b\u044c\u0448\u0435!<\/p>\n<pre><code class=\"python\">class CausalSelfAttetion(nn.Module):     def __init__(self, config):         super().__init__()         self.config = config         self.c_attn = nn.Linear(config.n_embd, 3 * config.n_embd, bias=True)         self.c_proj = nn.Linear(config.n_embd, config.n_embd, bias=True)      def forward(self, x):         B, T, C = x.shape         qkv = self.c_attn(x)         q, k, v = Tensor.split(qkv, 3, axis=2)         n_head = self.config.n_head         q = q.reshape(B, T, n_head, C \/\/ n_head).transpose(0, 2, 1, 3)         k = k.reshape(B, T, n_head, C \/\/ n_head).transpose(0, 2, 1, 3)         v = v.reshape(B, T, n_head, C \/\/ n_head).transpose(0, 2, 1, 3)         att = q @ k.transpose(0, -3, -1, -2) * (k.shape[-1] ** -0.5)         lg = att.local_gradients         att = Tensor.tril(att)         att[att == 0] = float('-inf')         att.local_gradients = lg         probs = Tensor.softmax(att, axis=-1)         y = probs @ v         y = y.transpose(0, 2, 1, 3).reshape(B, T, C)         y = self.c_proj(y)         return y<\/code><\/pre>\n<p>\u0421\u043b\u043e\u0439 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u044f, \u044f \u043f\u0440\u0435\u0434\u043f\u043e\u043b\u0430\u0433\u0430\u044e, \u0447\u0442\u043e \u0443\u0436\u0435 \u0432\u0438\u0434\u0435\u043b\u0438 \u043f\u043e\u0434\u043e\u0431\u043d\u0443\u044e \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044e, \u043f\u043e\u044d\u0442\u043e\u043c\u0443 \u043c\u043d\u043e\u0433\u043e \u0443\u0434\u0435\u043b\u044f\u0442\u044c \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u044f \u043d\u0435 \u0431\u0443\u0434\u0443. \u0412\u0438\u0434\u0438\u043c, \u043d\u0435\u0437\u043d\u0430\u043a\u043e\u043c\u044b\u0435 \u043d\u0430\u043c \u043c\u0435\u0442\u043e\u0434\u044b <code>Tensor.split(), Tensor.tril<\/code>, \u043e\u043d\u0438 \u0434\u0435\u043b\u0430\u044e\u0442 \u0442\u043e\u0436\u0435 \u0441\u0430\u043c\u043e\u0435, \u0447\u0442\u043e \u0438 \u0438\u0445 \u0430\u043d\u0430\u043b\u043e\u0433\u0438 \u0432 <code>pytorch<\/code><\/p>\n<pre><code class=\"python\">@classmethod def tril(cls, input, diagonal=0): value = np.tril(input.value, k=diagonal) local_gradients = ( ('tril', input, lambda x: x * np.tril(np.ones_like(input.value), k=diagonal)), ) return cls(value, local_gradients=local_gradients)  @classmethod def split(cls, array, split_size_or_sections, axis=0): value = np.split(array.value, split_size_or_sections, axis=axis) return cls(value, requires_grad=False)<\/code><\/pre>\n<p>\u041e\u0431\u0440\u0430\u0442\u0438\u0442\u0435 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u0435, \u044f \u043d\u0435 \u043e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u043b \u0432\u044b\u0447\u0438\u0441\u043b\u0435\u043d\u0438\u0435 \u0433\u0440\u0430\u0434\u0438\u0435\u043d\u0442\u0430 \u0434\u043b\u044f \u0432\u0442\u043e\u0440\u043e\u0439 \u043e\u043f\u0435\u0440\u0430\u0446\u0438\u0438. \u0417\u043d\u0430\u0447\u0438\u0442, \u043b\u0438\u0431\u043e \u043c\u043d\u0435 \u043f\u0440\u0438\u0434\u0451\u0442\u0441\u044f \u0435\u0433\u043e \u043e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u0442\u044c, \u043b\u0438\u0431\u043e \u044f \u043f\u0440\u043e\u0441\u0442\u043e \u043d\u0435 \u0441\u043c\u043e\u0433\u0443 \u043e\u0431\u0443\u0447\u0430\u0442\u044c \u043c\u043e\u0434\u0435\u043b\u0438!<\/p>\n<pre><code class=\"python\">class Block(nn.Module):     def __init__(self, config):         super().__init__()         self.ln_1 = nn.LayerNorm(config.n_embd)         self.attn = CausalSelfAttetion(config)         self.ln_2 = nn.LayerNorm(config.n_embd)         self.mlp = MLP(config)      def forward(self, x):         x = x + self.attn(self.ln_1(x))         x = x + self.mlp(self.ln_2(x))         return x<\/code><\/pre>\n<p>\u0422\u0443\u0442 \u044f \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044e \u0433\u043e\u0442\u043e\u0432\u044b\u0439 \u0441\u043b\u043e\u0439 <code>candle.nn.LayerNorm()<\/code>, \u0437\u0430\u0433\u043b\u044f\u043d\u0435\u043c \u0432 \u043d\u0435\u0433\u043e!<\/p>\n<pre><code class=\"python\">class LayerNorm:     def __init__(self, dim, eps=1e-5):         super().__init__()         self.param = None         self.gamma = Tensor.ones(dim, requires_grad=True)         self.beta = Tensor.zeros(dim, requires_grad=True)         self.eps = eps         self.all_layers = [self.gamma, self.beta]         self.grad = None      def __call__(self, x):         xmean = Tensor.mean(x, axis=2, keepdims=True)         xstd = Tensor.std(x, axis=2, keepdims=True)         x = (x - xmean) \/ (xstd + self.eps)         return self.gamma * x + self.beta<\/code><\/pre>\n<p>\u0412 \u0446\u0435\u043b\u043e\u043c \u043d\u0438\u0447\u0435\u0433\u043e \u0441\u043b\u043e\u0436\u043d\u043e\u0433\u043e, \u0435\u0441\u043b\u0438 \u0432\u044b \u0438\u0434\u0435\u0439\u043d\u043e \u0437\u043d\u0430\u043a\u043e\u043c\u044b \u0441 \u043c\u0435\u0442\u043e\u0434\u0430\u043c\u0438 \u043d\u043e\u0440\u043c\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u0438. \u0422\u0443\u0442 \u044f \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044e <code>Tensor.mean(), Tensor.std()<\/code>. \u041f\u043e\u0441\u043c\u043e\u0442\u0440\u0438\u043c \u043d\u0430 \u0438\u0445 \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044e<\/p>\n<pre><code class=\"python\">@classmethod def mean(cls, array, axis=None, keepdims=False): if axis == None: return Tensor.sum(array, axis=None, keepdims=keepdims) \/ np.size(array.value) else: delimeter = 1 if not isinstance(axis, int): for ax in axis: delimeter = delimeter * array.shape[ax] else: delimeter = array.shape[axis]  return Tensor.sum(array, axis=axis, keepdims=keepdims) \/ delimeter  @classmethod def std(cls, array, axis=None, keepdims=False):  if axis == None or axis == 0: mean = Tensor.mean(array, axis=axis, keepdims=False) sub = array - mean squared = sub ** 2 scaled_sum = Tensor.mean(squared, axis=axis, keepdims=keepdims) std = Tensor.sqrt(scaled_sum) return std  elif axis &gt;= 1: mean = Tensor.mean(array, axis=axis, keepdims=True) sub = array - mean squared = sub ** 2 scaled_sum = Tensor.mean(squared, axis=axis, keepdims=keepdims) std = Tensor.sqrt(scaled_sum) out = Tensor(std, local_gradients=None) out.local_gradients = (('std', std, lambda x: x * array.shape[axis] \/ (array.shape[axis] - 1)),) # array.shape[axis] \/ (array.shape[axis] - 1) additional multiplier due to dissimilarity return out <\/code><\/pre>\n<pre><code class=\"python\">class GPT(nn.Module):     def __init__(self, config):         super().__init__()         self.config = config         self.wte = nn.Embedding(config.vocab_size, config.n_embd)         self.wpe = nn.Embedding(1024, config.n_embd)         self.h = nn.ModuleList([Block(config) for _ in range(config.n_layer)])         self.ln_f = nn.LayerNorm(config.n_embd)         self.lm_head = nn.Linear(config.n_embd, config.vocab_size, bias=False)      def forward(self, x):         B, T = x.shape         assert T &gt;= self.config.block_size, f\"Cannot forward the sequence of length {T}, block_size is smaller\"         pos = Tensor.arange(T).reshape(1, -1)         pos_embd = self.wpe(pos)         tok_embd = self.wte(x)         x = pos_embd + tok_embd          for block in self.h:             x = block(x)         x = self.ln_f(x)         logits = self.lm_head(x)         return logits <\/code><\/pre>\n<p>\u0420\u0430\u0437\u0431\u0435\u0440\u0435\u043c\u0441\u044f \u043a\u0430\u043a \u0440\u0430\u0431\u043e\u0442\u0430\u044e\u0442 <code>candle.nn.Embedding() \u0438 candle.nn.ModuleList()<\/code>.<\/p>\n<pre><code class=\"python\">class Embedding:      def __init__(self, num_emb, emb_dim):         self.w = Tensor.randn((num_emb, emb_dim), requires_grad=True)         self.w.local_gradients = None         self.num_embd = num_emb         self.emb_dim = emb_dim         self.param = None         self.grad = None         self.all_layers = [self.w]      def __call__(self, x):          if self.param:             global Parameter             Parameter = self.param                      def multiply_by_locgrad(path_value):             temp = np.zeros_like(self.w.value)             np.add.at(np.zeros_like(self.w.value), x.value, path_value)             return temp                      x.value = x.value.astype(int)         local_gradients = (('embd', self.w, multiply_by_locgrad),)         return Tensor(self.w.value[x.value], local_gradients=local_gradients)<\/code><\/pre>\n<p>\u0412 \u0446\u0435\u043b\u043e\u043c \u043d\u0438\u0447\u0435\u0433\u043e \u0441\u043b\u043e\u0436\u043d\u043e\u0433\u043e, \u043f\u0440\u043e\u0441\u0442\u043e \u043e\u0436\u0438\u0434\u0430\u0435\u043c \u043d\u0430 \u0432\u0445\u043e\u0434\u0435 \u0442\u0435\u043d\u0437\u043e\u0440 \u0438\u0437 \u0446\u0435\u043b\u044b\u0445 \u0437\u043d\u0430\u0447\u0435\u043d\u0438\u0439 \u0438 \u0440\u0430\u0441\u0441\u043c\u0430\u0442\u0440\u0438\u0432\u0430\u0435\u043c \u044d\u0442\u0438 \u0437\u043d\u0430\u0447\u0435\u043d\u0438\u044f \u043a\u0430\u043a \u0438\u043d\u0434\u0435\u043a\u0441\u044b \u0434\u043b\u044f \u043c\u0430\u0442\u0440\u0438\u0446\u044b, \u0445\u0440\u0430\u043d\u044f\u0449\u0435\u0439 \u044d\u043c\u0431\u0435\u0434\u0434\u0438\u043d\u0433\u0438.<\/p>\n<pre><code class=\"python\">class ModuleList:      def __init__(self, layers):         self.layers = layers         self.index = 0      def __call__(self, x):         for layer in self.layers:             x = layer(x)         return x      def __len__(self):         return len(self.layers)      def __iter__(self):         self.index = 0         return self      def __next__(self):         if self.index &amp;lt; len(self.layers):             result = self.layers[self.index]             self.index += 1             return result         else:             raise StopIteration      def __getitem__(self, index):         return self.layers[index]<\/code><\/pre>\n<p>\u041e\u043a\u0430\u0437\u044b\u0432\u0430\u0435\u0442\u0441\u044f <code>ModuleList<\/code>\u044d\u0442\u043e \u043f\u0440\u043e\u0441\u0442\u043e \u0438\u0442\u0435\u0440\u0430\u0442\u043e\u0440!<br \/><strong>Note!<\/strong> \u0414\u043b\u044f \u043a\u043e\u0440\u0440\u0435\u043a\u0442\u043d\u043e\u0439 \u0440\u0430\u0431\u043e\u0442\u044b, \u043d\u0430\u043c \u043d\u0443\u0436\u043d\u043e \u043e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u0442\u044c \u043c\u0435\u0442\u043e\u0434 <code>Module.__setattr__()<\/code>, \u0438\u043d\u0430\u0447\u0435 \u0441\u043b\u043e\u0438 \u043a\u043e\u0442\u043e\u0440\u044b\u0435 \u043d\u0430\u0445\u043e\u0434\u044f\u0442\u0441\u044f \u0432\u043d\u0443\u0442\u0440\u0438 <code>ModuleList<\/code>, \u043f\u0440\u043e\u0441\u0442\u043e \u043d\u0435 \u0431\u0443\u0434\u0443\u0442 \u0432\u0438\u0434\u043d\u044b \u043d\u0430\u0448\u0435\u0439 \u0433\u043b\u043e\u0431\u0430\u043b\u044c\u043d\u043e\u0439 \u043f\u0435\u0440\u0435\u043c\u0435\u043d\u043d\u043e\u0439 <code>Parameter<\/code>, \u044f \u043d\u0435 \u0431\u0443\u0434\u0443 \u043d\u0430 \u044d\u0442\u043e\u043c \u043e\u0441\u0442\u0430\u043d\u0430\u0432\u043b\u0438\u0432\u0430\u0442\u044c\u0441\u044f, \u043d\u043e \u043f\u0440\u0438 \u0436\u0435\u043b\u0430\u043d\u0438\u0438 \u043c\u043e\u0436\u043d\u043e \u0437\u0430\u0433\u043b\u044f\u043d\u0443\u0442\u044c \u0432 \u043a\u043e\u0434 \u0438 \u0440\u0430\u0437\u043e\u0431\u0440\u0430\u0442\u044c\u0441\u044f!<\/p>\n<pre><code class=\"python\">class GPT(nn.Module):    @classmethod     def from_pretrained(cls, model_type):         assert model_type in ('gpt2', 'gpt2-medium', 'gpt2-large', 'gpt2-xl')         config_args = {             'gpt2': dict(n_layer=12, n_head=12, n_embd=768),             'gpt2-medium': dict(n_layer=24, n_head=16, n_embd=1024),             'gpt2-large': dict(n_layer=36, n_head=20, n_embd=1280),             'gpt2-xl': dict(n_layer=48, n_head=25, n_embd=1600)         }[model_type]         config_args['vocab_size'] = 50257         config_args['block_size'] = 1024          config = GPTConfig(**config_args)         model = GPT(config)         model = GPT.get_params(model, model_type)         return model<\/code><\/pre>\n<p>\u0417\u0434\u0435\u0441\u044c \u043c\u044b \u043e\u043f\u0440\u0435\u0434\u0435\u043b\u044f\u0435\u043c \u043c\u043e\u0434\u0435\u043b\u044c \u0438\u0437 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0438 <code>hf.transformers<\/code>, \u044d\u0442\u043e \u0441\u0430\u043c\u044b\u0439 \u043f\u0440\u043e\u0441\u0442\u043e\u0439 \u043f\u0443\u0442\u044c \u0434\u043b\u044f \u0434\u043e\u0441\u0442\u0443\u043f\u0430 \u043a \u0432\u0435\u0441\u0430\u043c GPT2.<\/p>\n<pre><code class=\"python\">@staticmethod def get_params(model, model_type):  from transformers import GPT2LMHeadModel, logging  logging.set_verbosity_error()  model_hf = GPT2LMHeadModel.from_pretrained(model_type) sd_hf = model_hf.state_dict() ...<\/code><\/pre>\n<p>\u042d\u0442\u043e \u0447\u0430\u0441\u0442\u044c \u043a\u043e\u0434\u0430, \u043e\u0442\u0432\u0435\u0447\u0430\u044e\u0449\u0435\u0433\u043e \u0437\u0430 \u043f\u0435\u0440\u0435\u043d\u043e\u0441 \u0432\u0435\u0441\u043e\u0432 \u0441 \u043c\u043e\u0434\u0435\u043b\u0438 <code>huggingface<\/code> \u043d\u0430 \u043d\u0430\u0448\u0443 \u043c\u043e\u0434\u0435\u043b\u044c. \u0422\u0430\u043c \u043d\u0438\u0447\u0435\u0433\u043e \u0443\u043c\u043d\u043e\u0433\u043e, \u043a\u0440\u043e\u043c\u0435 \u0430\u043a\u043a\u0443\u0440\u0430\u0442\u043d\u043e\u0439 \u0440\u0430\u0431\u043e\u0442\u044b.<\/p>\n<h2>\u0413\u0435\u043d\u0435\u0440\u0430\u0446\u0438\u044f<\/h2>\n<pre><code class=\"python\">model = GPT.from_pretrained(model_type=model_type)  max_length = 100 number_of_examples = 3 starting_sentence = \"Hello, I'm a language model,\"  enc = tiktoken.get_encoding('gpt2') tokens = enc.encode(starting_sentence) tokens = Tensor(tokens, dtype=int) tokens = Tensor.unsqueeze(tokens, 0) tokens = Tensor.repeat(tokens, number_of_examples, 0) x = tokens topk = 50 # responsible for \"creativity\" or \"adequacy\"  for i in tqdm(range(max_length)): logits = model(x) logits = logits[:, -1, :] probs = Tensor.softmax(logits, axis=-1) topk_indices, topk_probs = Tensor.topk(probs, topk) new_token = Tensor.multinomial_from_array(topk_indices, topk_probs, num_samples=1).reshape(-1, 1) x = Tensor.cat([x, new_token], axis=1)  for sample in x:     sample = sample.value.astype(int).tolist()     print(enc.decode(sample), end='\\n')     print('---------------') <\/code><\/pre>\n<pre><code>Hello, I'm a language model, not a coding model, but a compiler model.  I am thinking of what it means to be a developer and a language model  and to be able to write in some simple yet expressive way code that makes it  possible to use it with our software in any meaningful way to make  applications more readable, better, more readable. I can think of more words to talk about this. In this post, I'm trying to explain some of the main differences --------------- Hello, I'm a language model, based on my personal experience with code and development using C#.  This tutorial will help you create my own virtual language that you can use to implement things like an HTML page in your own app.  Now, what if I were going to learn how to program a website in an Objective-C program and create a web app out of it,  but didn't know how to do that and want to do some extra work to try it out? First I --------------- Hello, I'm a language model, I'm my own language. Well, a language model has two parts: a semantic construct and a structural construct. To understand the relationship with the semantic construct, let's see how all of the variables on a graph come together. Now I know that some of the variables on a graph represent numbers. But when you see the graphs, they're graphs. In other words, you see the connections and they are networks. We have an example<\/code><\/pre>\n<p>\u0412\u044b \u043c\u043e\u0436\u0435\u0442\u0435 \u043f\u043e\u0438\u0433\u0440\u0430\u0442\u044c\u0441\u044f \u0441 \u044d\u0442\u043e\u0439 \u043c\u043e\u0434\u0435\u043b\u044c\u044e \u043d\u0430 <a href=\"https:\/\/www.kaggle.com\/code\/freacle\/pretrained-gpt2-on-numpy\" rel=\"noopener noreferrer nofollow\">\u043a\u0430\u0433\u0433\u043b\u0435<\/a> \u0438\u043b\u0438 \u0441\u043a\u0430\u0447\u0430\u0432 \u043a \u0441\u0435\u0431\u0435 \u0438\u0437 \u0440\u0435\u043f\u043e\u0437\u0438\u0442\u043e\u0440\u0438\u044f <a href=\"https:\/\/github.com\/TimaGitHub\/smolGPT\" rel=\"noopener noreferrer nofollow\">smolGPT<\/a><\/p>\n<p>\u0412\u043e\u0442 \u0438 \u043f\u043e\u0434\u043e\u0448\u043b\u0430 \u043a \u043a\u043e\u043d\u0446\u0443 \u043c\u043e\u044f \u043c\u0438\u043d\u0438-\u0441\u0435\u0440\u0438\u044f \u043f\u043e \u0441\u043e\u0437\u0434\u0430\u043d\u0438\u044e \u0441\u043e\u0431\u0441\u0442\u0432\u0435\u043d\u043d\u043e\u0439 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0438 \u043d\u0430 <strong>NumPy<\/strong>.<br \/>\u0411\u043b\u0430\u0433\u043e\u0434\u0430\u0440\u044e \u0437\u0430 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u0435!<\/p>\n<p>\u041d\u0430\u0434\u0435\u044e\u0441\u044c, \u0432\u044b \u043d\u0430\u0448\u043b\u0438 \u0434\u043b\u044f \u0441\u0435\u0431\u044f \u044d\u0442\u043e\u0442 \u0446\u0438\u043a\u043b \u0441\u0442\u0430\u0442\u0435\u0439 \u043f\u043e\u0437\u043d\u0430\u0432\u0430\u0442\u0435\u043b\u044c\u043d\u044b\u043c!<\/p>\n<p>\u0415\u0441\u043b\u0438 \u0432\u0430\u043c \u043f\u043e\u043d\u0440\u0430\u0432\u0438\u043b\u043e\u0441\u044c, \u043f\u043e\u0436\u0430\u043b\u0443\u0439\u0441\u0442\u0430, \u043f\u043e\u0434\u0435\u043b\u0438\u0442\u0435\u0441\u044c \u0438\u043c\u0438, \u043f\u043e\u0441\u0442\u0430\u0432\u044c\u0442\u0435 upvote \u0438 \u043f\u043e\u0441\u0442\u0430\u0432\u044c\u0442\u0435 \u0437\u0432\u0451\u0437\u0434\u043e\u0447\u043a\u0438 \u043c\u043e\u0438\u043c \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044f\u043c \u043d\u0430 \u0433\u0438\u0442\u0445\u0430\u0431\u0435.<\/p>\n<p><a href=\"https:\/\/github.com\/TimaGitHub\/pycandle\" rel=\"noopener noreferrer nofollow\">\u041f\u0435\u0440\u0432\u0430\u044f \u0432\u0435\u0440\u0441\u0438\u044f \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0438<\/a><\/p>\n<p><a href=\"https:\/\/github.com\/TimaGitHub\/pycandle-2\" rel=\"noopener noreferrer nofollow\">\u0412\u0442\u043e\u0440\u0430\u044f \u0432\u0435\u0440\u0441\u0438\u044f \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0438<\/a><\/p>\n<p><a href=\"https:\/\/github.com\/TimaGitHub\/smolGPT\" rel=\"noopener noreferrer nofollow\">GPT-2 \u043d\u0430 \u044d\u0442\u043e\u0439 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0435<\/a><\/p>\n<\/div>\n<\/div>\n<\/div>\n<p><!----><!----><\/div>\n<p><!----><!----><br \/> \u0441\u0441\u044b\u043b\u043a\u0430 \u043d\u0430 \u043e\u0440\u0438\u0433\u0438\u043d\u0430\u043b \u0441\u0442\u0430\u0442\u044c\u0438 <a href=\"https:\/\/habr.com\/ru\/articles\/870504\/\"> https:\/\/habr.com\/ru\/articles\/870504\/<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<div><!--[--><!--]--><\/div>\n<div id=\"post-content-body\">\n<div>\n<div class=\"article-formatted-body article-formatted-body article-formatted-body_version-2\">\n<div xmlns=\"http:\/\/www.w3.org\/1999\/xhtml\">\n<p><strong>PyTorch<\/strong>\u00a0\u2014 \u044d\u0442\u043e \u043c\u043e\u0449\u043d\u044b\u0439 \u0438 \u0433\u0438\u0431\u043a\u0438\u0439 \u0444\u0440\u0435\u0439\u043c\u0432\u043e\u0440\u043a \u0434\u043b\u044f \u043c\u0430\u0448\u0438\u043d\u043d\u043e\u0433\u043e \u043e\u0431\u0443\u0447\u0435\u043d\u0438\u044f, \u0448\u0438\u0440\u043e\u043a\u043e \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u0435\u043c\u044b\u0439 \u0434\u043b\u044f \u0441\u043e\u0437\u0434\u0430\u043d\u0438\u044f \u043d\u0435\u0439\u0440\u043e\u043d\u043d\u044b\u0445 \u0441\u0435\u0442\u0435\u0439. \u041e\u043d \u043e\u0441\u043e\u0431\u0435\u043d\u043d\u043e \u043f\u043e\u043f\u0443\u043b\u044f\u0440\u0435\u043d \u0431\u043b\u0430\u0433\u043e\u0434\u0430\u0440\u044f \u043f\u0440\u043e\u0441\u0442\u043e\u0442\u0435 \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u043d\u0438\u044f, \u0434\u0438\u043d\u0430\u043c\u0438\u0447\u0435\u0441\u043a\u0438\u043c \u0432\u044b\u0447\u0438\u0441\u043b\u0438\u0442\u0435\u043b\u044c\u043d\u044b\u043c \u0433\u0440\u0430\u0444\u0430\u043c \u0438 \u0431\u043e\u0433\u0430\u0442\u043e\u0439 \u044d\u043a\u043e\u0441\u0438\u0441\u0442\u0435\u043c\u0435 \u0438\u043d\u0441\u0442\u0440\u0443\u043c\u0435\u043d\u0442\u043e\u0432 \u0434\u043b\u044f \u043e\u0431\u0443\u0447\u0435\u043d\u0438\u044f \u043c\u043e\u0434\u0435\u043b\u0435\u0439. \u0414\u043b\u044f \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u043d\u0438\u044f \u044d\u0442\u043e\u0433\u043e \u0444\u0440\u0435\u0439\u043c\u0432\u043e\u0440\u043a\u0430, \u0447\u0430\u0441\u0442\u043e \u0434\u043e\u0441\u0442\u0430\u0442\u043e\u0447\u043d\u043e \u043f\u043e\u0432\u0435\u0440\u0445\u043d\u043e\u0441\u0442\u043d\u043e \u043f\u043e\u043d\u0438\u043c\u0430\u0442\u044c \u0440\u0430\u0431\u043e\u0442\u0443 \u0430\u043b\u0433\u043e\u0440\u0438\u0442\u043c\u043e\u0432 \u043c\u0430\u0448\u0438\u043d\u043d\u043e\u0433\u043e \u043e\u0431\u0443\u0447\u0435\u043d\u0438\u044f.<\/p>\n<p>\u041d\u043e \u0410\u043d\u0434\u0440\u0435\u0439 \u041a\u0430\u0440\u043f\u0430\u0442\u044b, \u0438\u0437\u0432\u0435\u0441\u0442\u043d\u044b\u0439 \u0438\u0441\u0441\u043b\u0435\u0434\u043e\u0432\u0430\u0442\u0435\u043b\u044c \u0432 \u043e\u0431\u043b\u0430\u0441\u0442\u0438 \u0418\u0418, \u0441\u0447\u0438\u0442\u0430\u0435\u0442, \u0447\u0442\u043e \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044f \u0430\u043b\u0433\u043e\u0440\u0438\u0442\u043c\u043e\u0432 \u0441 \u043d\u0443\u043b\u044f \u043f\u043e\u0437\u0432\u043e\u043b\u044f\u0435\u0442 \u043f\u043e\u043d\u044f\u0442\u044c \u0438\u0445 \u0441\u0443\u0442\u044c \u0438 \u0434\u0435\u0442\u0430\u043b\u0438 \u0440\u0430\u0431\u043e\u0442\u044b, \u0447\u0442\u043e \u0441\u043b\u043e\u0436\u043d\u043e \u043e\u0441\u043e\u0437\u043d\u0430\u0442\u044c, \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044f \u0442\u043e\u043b\u044c\u043a\u043e \u0433\u043e\u0442\u043e\u0432\u044b\u0435 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0438. \u042d\u0442\u043e \u043f\u043e\u043c\u043e\u0433\u0430\u0435\u0442 \u0440\u0430\u0437\u0432\u0438\u0442\u044c \u0438\u043d\u0442\u0443\u0438\u0446\u0438\u044e \u0434\u043b\u044f \u0434\u0430\u043b\u044c\u043d\u0435\u0439\u0448\u0435\u0433\u043e \u043f\u0440\u0438\u043c\u0435\u043d\u0435\u043d\u0438\u044f \u0438 \u0443\u043b\u0443\u0447\u0448\u0435\u043d\u0438\u044f \u043c\u0435\u0442\u043e\u0434\u043e\u0432. \u0410\u043d\u0434\u0440\u0435\u0439 \u043f\u043e\u0441\u0432\u044f\u0449\u0430\u0435\u0442 \u043c\u043d\u043e\u0433\u043e \u0441\u043e\u0431\u0441\u0442\u0432\u0435\u043d\u043d\u043e\u0433\u043e \u0432\u0440\u0435\u043c\u0435\u043d\u0438, \u0447\u0442\u043e\u0431\u044b \u043e\u0431\u044a\u044f\u0441\u043d\u044f\u0442\u044c \u043a\u043b\u044e\u0447\u0435\u0432\u044b\u0435 \u043f\u0440\u0438\u043d\u0446\u0438\u043f\u044b \u0440\u0430\u0431\u043e\u0442\u044b \u043d\u0435\u0439\u0440\u043e\u0441\u0435\u0442\u0435\u0439 \u0432 \u0441\u0432\u043e\u0438\u0445\u00a0<a href=\"https:\/\/karpathy.ai\/\" rel=\"noopener noreferrer nofollow\">\u0431\u043b\u043e\u0433\u0430\u0445<\/a>\u00a0\u0438 \u043d\u0430 \u0441\u0432\u043e\u0451\u043c\u00a0<a href=\"https:\/\/www.youtube.com\/@AndrejKarpathy\" rel=\"noopener noreferrer nofollow\">\u044e\u0442\u0443\u0431-\u043a\u0430\u043d\u0430\u043b\u0435<\/a>. \u041e\u043d \u0442\u0430\u043a\u0436\u0435 \u043d\u0435 \u0440\u0430\u0437 \u043f\u043e\u0434\u0447\u0435\u0440\u043a\u0438\u0432\u0430\u043b, \u0447\u0442\u043e \u043d\u0430 \u0435\u0433\u043e \u043a\u0443\u0440\u0441\u0435 \u0432\u00a0<a href=\"https:\/\/cs.stanford.edu\/people\/karpathy\/\" rel=\"noopener noreferrer nofollow\">C\u0442\u044d\u043d\u0444\u043e\u0440\u0434\u0435<\/a>\u00a0\u0435\u0441\u0442\u044c \u0437\u0430\u0434\u0430\u0447\u0438 \u043f\u043e \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u0438 \u0440\u0430\u0437\u043b\u0438\u0447\u043d\u044b\u0445 \u0430\u043b\u0433\u043e\u0440\u0438\u0442\u043c\u043e\u0432, \u043d\u0430\u043f\u0440\u0438\u043c\u0435\u0440, \u043e\u0431\u0440\u0430\u0442\u043d\u043e\u0435 \u0440\u0430\u0441\u043f\u0440\u043e\u0441\u0442\u0440\u0430\u043d\u0435\u043d\u0438\u0435.<\/p>\n<p>\u042f \u0445\u043e\u0442\u0435\u043b \u0431\u044b \u043f\u043e\u0441\u0432\u044f\u0442\u0438\u0442\u044c \u0434\u0430\u043d\u043d\u0443\u044e \u0441\u0442\u0430\u0442\u044c\u044e \u044d\u0442\u043e\u0439 \u0438\u0434\u0435\u0438, \u043f\u043e\u0442\u043e\u043c\u0443 \u0447\u0442\u043e \u043c\u043d\u0435 \u0441\u0430\u043c\u043e\u043c\u0443 \u043e\u0441\u043e\u0431\u0435\u043d\u043d\u043e \u0438\u043d\u0442\u0435\u0440\u0435\u0441\u043d\u043e \u043a\u043e\u043f\u0430\u0442\u044c\u0441\u044f \u0432 \u0430\u043b\u0433\u043e\u0440\u0438\u0442\u043c\u0430\u0445 \u0433\u043b\u0443\u0431\u043e\u043a\u043e\u0433\u043e \u043e\u0431\u0443\u0447\u0435\u043d\u0438\u044f. \u042d\u0442\u0430 \u0441\u0442\u0430\u0442\u044c\u044f \u043f\u0440\u043e\u0434\u043e\u043b\u0436\u0435\u043d\u0438\u0435\u00a0<a href=\"https:\/\/habr.com\/ru\/articles\/870426\/\" rel=\"noopener noreferrer nofollow\">\u0442\u0440\u0435\u0442\u044c\u0435\u0439 \u0441\u0442\u0430\u0442\u044c\u0438<\/a><\/p>\n<p>\u0418\u0442\u0430\u043a, \u043a\u0430\u043a \u044f \u0438 \u0441\u043a\u0430\u0437\u0430\u043b \u0432 \u043f\u0440\u0435\u0434\u044b\u0434\u0443\u0449\u0435\u0439 \u0441\u0442\u0430\u0442\u044c\u0435, \u0443 \u043d\u0430\u0441 \u0434\u043e\u0441\u0442\u0430\u0442\u043e\u0447\u043d\u043e \u0437\u043d\u0430\u043d\u0438\u0439, \u0447\u0442\u043e\u0431\u044b \u0441\u043e\u0431\u0440\u0430\u0442\u044c \u0438\u0445 \u0432 \u0446\u0435\u043b\u0443\u044e \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0443. \u0422\u0430\u043a \u044f \u0438 \u0441\u0434\u0435\u043b\u0430\u043b, \u0440\u0435\u0430\u043b\u0438\u0437\u043e\u0432\u0430\u0432 <a href=\"https:\/\/github.com\/TimaGitHub\/pycandle-2\" rel=\"noopener noreferrer nofollow\">pycandle<\/a>.<br \/>\u0422\u0430\u043c \u043d\u0430\u0445\u043e\u0434\u0438\u0442\u044c\u0441\u044f \u0432\u0441\u0451 \u0442\u043e, \u0447\u0442\u043e \u043c\u044b \u0438\u0437\u0443\u0447\u0430\u043b\u0438 \u043d\u0430 \u043f\u0440\u043e\u0442\u044f\u0436\u0435\u043d\u0438\u0438 3 \u0447\u0430\u0441\u0442\u0435\u0439. \u0412 \u044d\u0442\u043e\u0439 \u0447\u0430\u0441\u0442\u0438 \u043c\u044b \u0431\u0443\u0434\u0435\u043c \u043f\u0438\u0441\u0430\u0442\u044c \u0438\u043d\u0444\u0435\u0440\u0435\u043d\u0441 \u043a\u043e\u0434 \u0434\u043b\u044f <strong>GPT2<\/strong> \u043d\u0430 \u0441\u043e\u0431\u0441\u0442\u0432\u0435\u043d\u043d\u043e\u0439 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a\u0435!<br \/>\u041d\u0430\u0447\u043d\u0451\u043c \u0441 \u0438\u043c\u043f\u043e\u0440\u0442\u0430 \u0431\u0438\u0431\u043b\u0438\u043e\u0442\u0435\u043a<\/p>\n<pre><code class=\"python\">import torch import torch.nn as nn<\/code><\/pre>\n<p>\u041e\u0439, \u043d\u0435 \u0442\u043e. \u042f \u0438\u043c\u0435\u043b \u0432 \u0432\u0438\u0434\u0443<\/p>\n<pre><code class=\"python\">import candle import candle.nn as nn<\/code><\/pre>\n<p>\u0422\u0430\u043a\u0436\u0435 \u0432\u043e\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u0435\u043c\u0441\u044f \u0442\u043e\u043a\u0435\u043d\u0438\u0437\u0430\u0442\u043e\u0440\u043e\u043c \u043e\u0442 OpenAI, \u043a\u043e\u0442\u043e\u0440\u044b\u0439 \u043e\u043d\u0438 \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044e\u0442 \u0432 \u0441\u0432\u043e\u0438\u0445 \u043c\u043e\u0434\u0435\u043b\u044f\u0445. \u0421\u0434\u0435\u043b\u0430\u0435\u043c \u0432\u0441\u0435 \u043d\u0435\u043e\u0431\u0445\u043e\u0434\u0438\u043c\u044b\u0435 \u0438\u043c\u043f\u043e\u0440\u0442\u044b<\/p>\n<pre><code class=\"python\">import tiktoken from candle import Tensor from dataclasses import dataclass<\/code><\/pre>\n<p>\u041e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u043c \u043a\u043e\u043d\u0444\u0438\u0433\u0443\u0440\u0430\u0446\u0438\u044e \u043d\u0430\u0448\u0435\u0439 \u043c\u043e\u0434\u0435\u043b\u0438 \u0441 \u043f\u043e\u043c\u043e\u0449\u044c\u044e <code>dataclasses<\/code><\/p>\n<pre><code class=\"python\">@dataclass class GPTConfig:     block_size: int = 1024 # \u0440\u0430\u0437\u043c\u0435\u0440 \u043e\u043a\u043d\u0430 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u04d9     vocab_size: int = 50257 # \u0440\u0430\u0437\u043c\u0435\u0440 \u0441\u043b\u043e\u0432\u0430\u0440\u044f BPE     n_layer: int = 12 # \u043a\u043e\u043b\u0438\u0447\u0435\u0441\u0442\u0432\u043e \u0441\u043b\u043e\u0451\u0432     n_head: int = 12 # \u043a\u043e\u043b\u0438\u0447\u0435\u0441\u0442\u0432\u043e \u0433\u043e\u043b\u043e\u0432 \u0432 \u043c\u0435\u0445\u0430\u043d\u0438\u0437\u043c\u0435 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u044f     n_embd: int = 768 # \u0440\u0430\u0437\u043c\u0435\u0440\u043d\u043e\u0441\u0442\u044c \u0432\u0435\u043a\u0442\u043e\u0440\u0430 \u044d\u043c\u0431\u0435\u0434\u0434\u0438\u043d\u0433\u043e\u0432<\/code><\/pre>\n<p>\u041e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u043c \u043b\u0438\u043d\u0435\u0439\u043d\u044b\u0439 \u0441\u043b\u043e\u0439 \u043c\u043e\u0434\u0435\u043b\u0438<\/p>\n<pre><code class=\"python\">class MLP(nn.Module):     def __init__(self, config):         super().__init__()         self.c_fc = nn.Linear(config.n_embd, 4 * config.n_embd)         self.gelu = nn.GeLU()         self.c_proj = nn.Linear(4 * config.n_embd, config.n_embd)      def forward(self, x):         x = self.c_fc(x)         x = self.gelu(x)         return self.c_proj(x)<\/code><\/pre>\n<p><strong>\u041e\u0431\u0440\u0430\u0442\u0438\u0442\u0435 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u0435<\/strong> \u0442\u0443\u0442 <code>candle.nn<\/code>, \u0430 <strong>\u043d\u0435<\/strong> <code>torch.nn<\/code>. \u0421 \u043b\u0438\u043d\u0435\u0439\u043d\u044b\u043c\u0438 \u0441\u043b\u043e\u044f\u043c\u0438 \u043c\u044b \u0443\u0436\u0435 \u0437\u043d\u0430\u043a\u043e\u043c\u044b, \u0430 <code>GeLU()<\/code>, \u044d\u0442\u043e \u0447\u0442\u043e \u0442\u0430\u043a\u043e\u0435? <\/p>\n<figure class=\"full-width\"><\/figure>\n<p>\u0418\u0441\u0441\u043b\u0435\u0434\u043e\u0432\u0430\u0442\u0435\u043b\u0438 \u0438\u0437 <code>OpenAI<\/code>, \u0432\u044b\u044f\u0441\u043d\u0438\u043b\u0438, \u0447\u0442\u043e \u0432 \u0438\u0445 \u043c\u043e\u0434\u0435\u043b\u044f\u0445 \u0442\u0430\u043a\u0430\u044f \u0444\u0443\u043d\u043a\u0446\u0438\u044f \u0440\u0430\u0431\u043e\u0442\u0430\u0435\u0442 \u043b\u0443\u0447\u0448\u0435 \u0432\u0441\u0435\u0433\u043e, \u043f\u043e\u0441\u043c\u043e\u0442\u0440\u0438\u043c \u043d\u0430 \u0435\u0451 \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044e<\/p>\n<figure class=\"full-width\"><\/figure>\n<pre><code class=\"python\">class GeLU:     def __init__(self):         Parameter([self,[]])      def __call__(self, x):         return Tensor.gelu(x)         #return 0.5 * x * (1 + Tensor.tanh(0.79788456 * (x + 0.044715 * (x ** 3))))<\/code><\/pre>\n<p>\u042f \u0437\u0430\u043a\u043e\u043c\u043c\u0435\u043d\u0442\u0438\u0440\u043e\u0432\u0430\u043b \u0432\u0442\u043e\u0440\u0443\u044e \u0441\u0442\u0440\u043e\u0447\u043a\u0443 \u0438 \u0432\u043c\u0435\u0441\u0442\u043e \u044d\u0442\u043e\u0433\u043e \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044e \u0432\u0441\u0442\u0440\u043e\u0435\u043d\u043d\u044b\u0439 \u043c\u0435\u0442\u043e\u0434 <code>Tensor.gelu()<\/code>, \u043d\u043e \u043d\u0438\u043a\u0442\u043e \u043d\u0435 \u0437\u0430\u043f\u0440\u0435\u0449\u0430\u0435\u0442 \u0438 \u0447\u0435\u0441\u0442\u043d\u043e \u0441\u0447\u0438\u0442\u0430\u0442\u044c \u043f\u043e \u0444\u043e\u0440\u043c\u0443\u043b\u0435. \u0417\u0430\u0433\u043b\u044f\u043d\u0435\u043c \u0432 \u044d\u0442\u043e\u0442 \u043c\u0435\u0442\u043e\u0434<\/p>\n<pre><code class=\"python\">from scipy.stats import norm @classmethod def gelu(cls, x): return x * norm.cdf(x.value) # \u0444\u043e\u0440\u043c\u0443\u043b\u0430 \u0438\u0437 \u043a\u0430\u0440\u0442\u0438\u043d\u043a\u0438<\/code><\/pre>\n<p>\u0422\u0435\u043f\u0435\u0440\u044c \u0432 \u043b\u0438\u043d\u0435\u0439\u043d\u043e\u043c \u0441\u043b\u043e\u0435 \u043d\u0430\u043c \u0432\u0441\u0435 \u043f\u043e\u043d\u044f\u0442\u043d\u043e, \u0438\u0434\u0451\u043c \u0434\u0430\u043b\u044c\u0448\u0435!<\/p>\n<pre><code class=\"python\">class CausalSelfAttetion(nn.Module):     def __init__(self, config):         super().__init__()         self.config = config         self.c_attn = nn.Linear(config.n_embd, 3 * config.n_embd, bias=True)         self.c_proj = nn.Linear(config.n_embd, config.n_embd, bias=True)      def forward(self, x):         B, T, C = x.shape         qkv = self.c_attn(x)         q, k, v = Tensor.split(qkv, 3, axis=2)         n_head = self.config.n_head         q = q.reshape(B, T, n_head, C \/\/ n_head).transpose(0, 2, 1, 3)         k = k.reshape(B, T, n_head, C \/\/ n_head).transpose(0, 2, 1, 3)         v = v.reshape(B, T, n_head, C \/\/ n_head).transpose(0, 2, 1, 3)         att = q @ k.transpose(0, -3, -1, -2) * (k.shape[-1] ** -0.5)         lg = att.local_gradients         att = Tensor.tril(att)         att[att == 0] = float('-inf')         att.local_gradients = lg         probs = Tensor.softmax(att, axis=-1)         y = probs @ v         y = y.transpose(0, 2, 1, 3).reshape(B, T, C)         y = self.c_proj(y)         return y<\/code><\/pre>\n<p>\u0421\u043b\u043e\u0439 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u044f, \u044f \u043f\u0440\u0435\u0434\u043f\u043e\u043b\u0430\u0433\u0430\u044e, \u0447\u0442\u043e \u0443\u0436\u0435 \u0432\u0438\u0434\u0435\u043b\u0438 \u043f\u043e\u0434\u043e\u0431\u043d\u0443\u044e \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044e, \u043f\u043e\u044d\u0442\u043e\u043c\u0443 \u043c\u043d\u043e\u0433\u043e \u0443\u0434\u0435\u043b\u044f\u0442\u044c \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u044f \u043d\u0435 \u0431\u0443\u0434\u0443. \u0412\u0438\u0434\u0438\u043c, \u043d\u0435\u0437\u043d\u0430\u043a\u043e\u043c\u044b\u0435 \u043d\u0430\u043c \u043c\u0435\u0442\u043e\u0434\u044b <code>Tensor.split(), Tensor.tril<\/code>, \u043e\u043d\u0438 \u0434\u0435\u043b\u0430\u044e\u0442 \u0442\u043e\u0436\u0435 \u0441\u0430\u043c\u043e\u0435, \u0447\u0442\u043e \u0438 \u0438\u0445 \u0430\u043d\u0430\u043b\u043e\u0433\u0438 \u0432 <code>pytorch<\/code><\/p>\n<pre><code class=\"python\">@classmethod def tril(cls, input, diagonal=0): value = np.tril(input.value, k=diagonal) local_gradients = ( ('tril', input, lambda x: x * np.tril(np.ones_like(input.value), k=diagonal)), ) return cls(value, local_gradients=local_gradients)  @classmethod def split(cls, array, split_size_or_sections, axis=0): value = np.split(array.value, split_size_or_sections, axis=axis) return cls(value, requires_grad=False)<\/code><\/pre>\n<p>\u041e\u0431\u0440\u0430\u0442\u0438\u0442\u0435 \u0432\u043d\u0438\u043c\u0430\u043d\u0438\u0435, \u044f \u043d\u0435 \u043e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u043b \u0432\u044b\u0447\u0438\u0441\u043b\u0435\u043d\u0438\u0435 \u0433\u0440\u0430\u0434\u0438\u0435\u043d\u0442\u0430 \u0434\u043b\u044f \u0432\u0442\u043e\u0440\u043e\u0439 \u043e\u043f\u0435\u0440\u0430\u0446\u0438\u0438. \u0417\u043d\u0430\u0447\u0438\u0442, \u043b\u0438\u0431\u043e \u043c\u043d\u0435 \u043f\u0440\u0438\u0434\u0451\u0442\u0441\u044f \u0435\u0433\u043e \u043e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u0442\u044c, \u043b\u0438\u0431\u043e \u044f \u043f\u0440\u043e\u0441\u0442\u043e \u043d\u0435 \u0441\u043c\u043e\u0433\u0443 \u043e\u0431\u0443\u0447\u0430\u0442\u044c \u043c\u043e\u0434\u0435\u043b\u0438!<\/p>\n<pre><code class=\"python\">class Block(nn.Module):     def __init__(self, config):         super().__init__()         self.ln_1 = nn.LayerNorm(config.n_embd)         self.attn = CausalSelfAttetion(config)         self.ln_2 = nn.LayerNorm(config.n_embd)         self.mlp = MLP(config)      def forward(self, x):         x = x + self.attn(self.ln_1(x))         x = x + self.mlp(self.ln_2(x))         return x<\/code><\/pre>\n<p>\u0422\u0443\u0442 \u044f \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044e \u0433\u043e\u0442\u043e\u0432\u044b\u0439 \u0441\u043b\u043e\u0439 <code>candle.nn.LayerNorm()<\/code>, \u0437\u0430\u0433\u043b\u044f\u043d\u0435\u043c \u0432 \u043d\u0435\u0433\u043e!<\/p>\n<pre><code class=\"python\">class LayerNorm:     def __init__(self, dim, eps=1e-5):         super().__init__()         self.param = None         self.gamma = Tensor.ones(dim, requires_grad=True)         self.beta = Tensor.zeros(dim, requires_grad=True)         self.eps = eps         self.all_layers = [self.gamma, self.beta]         self.grad = None      def __call__(self, x):         xmean = Tensor.mean(x, axis=2, keepdims=True)         xstd = Tensor.std(x, axis=2, keepdims=True)         x = (x - xmean) \/ (xstd + self.eps)         return self.gamma * x + self.beta<\/code><\/pre>\n<p>\u0412 \u0446\u0435\u043b\u043e\u043c \u043d\u0438\u0447\u0435\u0433\u043e \u0441\u043b\u043e\u0436\u043d\u043e\u0433\u043e, \u0435\u0441\u043b\u0438 \u0432\u044b \u0438\u0434\u0435\u0439\u043d\u043e \u0437\u043d\u0430\u043a\u043e\u043c\u044b \u0441 \u043c\u0435\u0442\u043e\u0434\u0430\u043c\u0438 \u043d\u043e\u0440\u043c\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u0438. \u0422\u0443\u0442 \u044f \u0438\u0441\u043f\u043e\u043b\u044c\u0437\u0443\u044e <code>Tensor.mean(), Tensor.std()<\/code>. \u041f\u043e\u0441\u043c\u043e\u0442\u0440\u0438\u043c \u043d\u0430 \u0438\u0445 \u0440\u0435\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044e<\/p>\n<pre><code class=\"python\">@classmethod def mean(cls, array, axis=None, keepdims=False): if axis == None: return Tensor.sum(array, axis=None, keepdims=keepdims) \/ np.size(array.value) else: delimeter = 1 if not isinstance(axis, int): for ax in axis: delimeter = delimeter * array.shape[ax] else: delimeter = array.shape[axis]  return Tensor.sum(array, axis=axis, keepdims=keepdims) \/ delimeter  @classmethod def std(cls, array, axis=None, keepdims=False):  if axis == None or axis == 0: mean = Tensor.mean(array, axis=axis, keepdims=False) sub = array - mean squared = sub ** 2 scaled_sum = Tensor.mean(squared, axis=axis, keepdims=keepdims) std = Tensor.sqrt(scaled_sum) return std  elif axis &gt;= 1: mean = Tensor.mean(array, axis=axis, keepdims=True) sub = array - mean squared = sub ** 2 scaled_sum = Tensor.mean(squared, axis=axis, keepdims=keepdims) std = Tensor.sqrt(scaled_sum) out = Tensor(std, local_gradients=None) out.local_gradients = (('std', std, lambda x: x * array.shape[axis] \/ (array.shape[axis] - 1)),) # array.shape[axis] \/ (array.shape[axis] - 1) additional multiplier due to dissimilarity return out <\/code><\/pre>\n<pre><code class=\"python\">class GPT(nn.Module):     def __init__(self, config):         super().__init__()         self.config = config         self.wte = nn.Embedding(config.vocab_size, config.n_embd)         self.wpe = nn.Embedding(1024, config.n_embd)         self.h = nn.ModuleList([Block(config) for _ in range(config.n_layer)])         self.ln_f = nn.LayerNorm(config.n_embd)         self.lm_head = nn.Linear(config.n_embd, config.vocab_size, bias=False)      def forward(self, x):         B, T = x.shape         assert T &gt;= self.config.block_size, f\"Cannot forward the sequence of length {T}, block_size is smaller\"         pos = Tensor.arange(T).reshape(1, -1)         pos_embd = self.wpe(pos)         tok_embd = self.wte(x)         x = pos_embd + tok_embd          for block in self.h:             x = block(x)         x = self.ln_f(x)         logits = self.lm_head(x)         return logits <\/code><\/pre>\n<p>\u0420\u0430\u0437\u0431\u0435\u0440\u0435\u043c\u0441\u044f \u043a\u0430\u043a \u0440\u0430\u0431\u043e\u0442\u0430\u044e\u0442 <code>candle.nn.Embedding() \u0438 candle.nn.ModuleList()<\/code>.<\/p>\n<pre><code class=\"python\">class Embedding:      def __init__(self, num_emb, emb_dim):         self.w = Tensor.randn((num_emb, emb_dim), requires_grad=True)         self.w.local_gradients = None         self.num_embd = num_emb         self.emb_dim = emb_dim         self.param = None         self.grad = None         self.all_layers = [self.w]      def __call__(self, x):          if self.param:             global Parameter             Parameter = self.param                      def multiply_by_locgrad(path_value):             temp = np.zeros_like(self.w.value)             np.add.at(np.zeros_like(self.w.value), x.value, path_value)             return temp                      x.value = x.value.astype(int)         local_gradients = (('embd', self.w, multiply_by_locgrad),)         return Tensor(self.w.value[x.value], local_gradients=local_gradients)<\/code><\/pre>\n<p>\u0412 \u0446\u0435\u043b\u043e\u043c \u043d\u0438\u0447\u0435\u0433\u043e \u0441\u043b\u043e\u0436\u043d\u043e\u0433\u043e, \u043f\u0440\u043e\u0441\u0442\u043e \u043e\u0436\u0438\u0434\u0430\u0435\u043c \u043d\u0430 \u0432\u0445\u043e\u0434\u0435 \u0442\u0435\u043d\u0437\u043e\u0440 \u0438\u0437 \u0446\u0435\u043b\u044b\u0445 \u0437\u043d\u0430\u0447\u0435\u043d\u0438\u0439 \u0438 \u0440\u0430\u0441\u0441\u043c\u0430\u0442\u0440\u0438\u0432\u0430\u0435\u043c \u044d\u0442\u0438 \u0437\u043d\u0430\u0447\u0435\u043d\u0438\u044f \u043a\u0430\u043a \u0438\u043d\u0434\u0435\u043a\u0441\u044b \u0434\u043b\u044f \u043c\u0430\u0442\u0440\u0438\u0446\u044b, \u0445\u0440\u0430\u043d\u044f\u0449\u0435\u0439 \u044d\u043c\u0431\u0435\u0434\u0434\u0438\u043d\u0433\u0438.<\/p>\n<pre><code class=\"python\">class ModuleList:      def __init__(self, layers):         self.layers = layers         self.index = 0      def __call__(self, x):         for layer in self.layers:             x = layer(x)         return x      def __len__(self):         return len(self.layers)      def __iter__(self):         self.index = 0         return self      def __next__(self):         if self.index &amp;lt; len(self.layers):             result = self.layers[self.index]             self.index += 1             return result         else:             raise StopIteration      def __getitem__(self, index):         return self.layers[index]<\/code><\/pre>\n<p>\u041e\u043a\u0430\u0437\u044b\u0432\u0430\u0435\u0442\u0441\u044f <code>ModuleList<\/code>\u044d\u0442\u043e \u043f\u0440\u043e\u0441\u0442\u043e \u0438\u0442\u0435\u0440\u0430\u0442\u043e\u0440!<br \/><strong>Note!<\/strong> \u0414\u043b\u044f \u043a\u043e\u0440\u0440\u0435\u043a\u0442\u043d\u043e\u0439 \u0440\u0430\u0431\u043e\u0442\u044b, \u043d\u0430\u043c \u043d\u0443\u0436\u043d\u043e \u043e\u043f\u0440\u0435\u0434\u0435\u043b\u0438\u0442\u044c \u043c\u0435\u0442\u043e\u0434 <code>Module.__setattr__()<\/code>, \u0438\u043d\u0430\u0447\u0435 \u0441\u043b\u043e\u0438 \u043a\u043e\u0442\u043e\u0440\u044b\u0435 \u043d\u0430\u0445\u043e\u0434\u044f\u0442\u0441\u044f \u0432\u043d\u0443\u0442\u0440\u0438 <code>ModuleList<\/code>, \u043f\u0440\u043e\u0441\u0442\u043e \u043d\u0435 \u0431\u0443\u0434\u0443\u0442 \u0432\u0438\u0434\u043d\u044b \u043d\u0430\u0448\u0435\u0439 \u0433\u043b\u043e\u0431\u0430\u043b\u044c\u043d\u043e\u0439 \u043f\u0435\u0440\u0435\u043c\u0435\u043d\u043d\u043e\u0439 <code>Parameter<\/code>, \u044f \u043d\u0435 \u0431\u0443\u0434\u0443 \u043d\u0430 \u044d\u0442\u043e\u043c<\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[],"tags":[],"class_list":["post-443737","post","type-post","status-publish","format-standard","hentry"],"_links":{"self":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts\/443737","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=443737"}],"version-history":[{"count":0,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts\/443737\/revisions"}],"wp:attachment":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=443737"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=443737"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=443737"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}