cpython

Commit Graph

Author	SHA1	Message	Date
Pablo Galindo Salgado	417206a05c	gh-99891: Fix infinite recursion in the tokenizer when showing warnings (GH-99893) Automerge-Triggered-By: GH:pablogsal	2022-11-30 03:36:06 -08:00
Pablo Galindo Salgado	e13d1d9dda	gh-99581: Fix a buffer overflow in the tokenizer when copying lines that fill the available buffer (#99605 )	2022-11-20 20:20:03 +00:00
Victor Stinner	4ce2a202c7	gh-99300: Use Py_NewRef() in Parser/ directory (#99330 ) Replace Py_INCREF() with Py_NewRef() in C files of the Parser/ directory and in the PEG generator.	2022-11-10 15:30:05 +01:00
Lysandros Nikolaou	3de08ce8c1	gh-97997: Add col_offset field to tokenizer and use that for AST nodes (#98000 )	2022-10-07 14:38:35 -07:00
Lysandros Nikolaou	cbf0afd8a1	gh-97973: Return all necessary information from the tokenizer (GH-97984) Right now, the tokenizer only returns type and two pointers to the start and end of the token. This PR modifies the tokenizer to return the type and set all of the necessary information, so that the parser does not have to this.	2022-10-06 16:07:17 -07:00
Pablo Galindo Salgado	aab01e3524	gh-96670: Raise SyntaxError when parsing NULL bytes (#97594 )	2022-09-27 23:23:42 +01:00
Matthias Görgens	81e36f350b	gh-96678: Fix UB of null pointer arithmetic (GH-96782) Automerge-Triggered-By: GH:pablogsal	2022-09-13 06:14:35 -07:00
Michael Droettboom	8bc356a7dd	gh-96268: Fix loading invalid UTF-8 (#96270 ) This makes tokenizer.c:valid_utf8 match stringlib/codecs.h:decode_utf8. It also fixes an off-by-one error introduced in 3.10 for the line number when the tokenizer reports bad UTF8.	2022-09-07 14:23:54 -07:00
Michael Droettboom	05692c67c5	gh-96611: Fix error message for invalid UTF-8 in mid-multiline string (#96623 )	2022-09-07 00:12:16 +01:00
Pablo Galindo Salgado	36fcde61ba	gh-94360: Fix a tokenizer crash when reading encoded files with syntax errors from stdin (#94386 ) * gh-94360: Fix a tokenizer crash when reading encoded files with syntax errors from stdin Signed-off-by: Pablo Galindo <pablogsal@gmail.com> * nitty nit Co-authored-by: Łukasz Langa <lukasz@langa.pl>	2022-07-05 17:39:21 +01:00
Serhiy Storchaka	6fd4c8ec77	gh-93741: Add private C API _PyImport_GetModuleAttrString() (GH-93742) It combines PyImport_ImportModule() and PyObject_GetAttrString() and saves 4-6 lines of code on every use. Add also _PyImport_GetModuleAttr() which takes Python strings as arguments.	2022-06-14 07:15:26 +03:00
Kumar Aditya	cb04a09d2d	GH-93207: Remove HAVE_STDARG_PROTOTYPES configure check for stdarg.h (#93215 )	2022-05-27 13:30:45 +02:00
Victor Stinner	5115a16831	gh-93103: Parser uses PyConfig.parser_debug instead of Py_DebugFlag (#93106 ) * Replace deprecated Py_DebugFlag with PyConfig.parser_debug in the parser. * Add Parser.debug member. * Add tok_state.debug member. * Py_FrozenMain(): Replace Py_VerboseFlag with PyConfig.verbose.	2022-05-24 22:35:08 +02:00
Victor Stinner	da5727a120	gh-92651: Remove the Include/token.h header file (#92652 ) Remove the token.h header file. There was never any public tokenizer C API. The token.h header file was only designed to be used by Python internals. Move Include/token.h to Include/internal/pycore_token.h. Including this header file now requires that the Py_BUILD_CORE macro is defined. It no longer checks for the Py_LIMITED_API macro. Rename functions: * PyToken_OneChar() => _PyToken_OneChar() * PyToken_TwoChars() => _PyToken_TwoChars() * PyToken_ThreeChars() => _PyToken_ThreeChars()	2022-05-11 23:22:50 +02:00
Serhiy Storchaka	43a8bf1ea4	gh-87999: Change warning type for numeric literal followed by keyword (GH-91980) The warning emitted by the Python parser for a numeric literal immediately followed by keyword has been changed from deprecation warning to syntax warning.	2022-04-27 20:15:14 +03:00
Christian Heimes	3df0e63aab	bpo-46315: Use fopencookie only on Emscripten 3.x and newer (GH-32266)	2022-04-02 23:11:38 +02:00
Hugo van Kemenade	6881ea936e	bpo-47126: Update to canonical PEP URLs specified by PEP 676 (GH-32124)	2022-03-30 12:00:27 +01:00
Christian Heimes	9b889b5bda	bpo-46315: Use fopencookie() to avoid dup() in _PyTokenizer_FindEncodingFilename (GH-32033) WASI does not have dup() and Emscripten's emulation is slow.	2022-03-22 17:08:51 +01:00
Oleg Iarygin	13b0412223	bpo-46920: Remove code that has explainers why it was disabled (GH-31813)	2022-03-14 17:04:22 +01:00
Serhiy Storchaka	090e5c4b94	bpo-46820: Fix a SyntaxError in a numeric literal followed by "not in" (GH-31479) Fix parsing a numeric literal immediately (without spaces) followed by "not in" keywords, like in "1not in x". Now the parser only emits a warning, not a syntax error.	2022-02-22 09:51:51 +02:00
Eric Snow	81c72044a1	bpo-46541: Replace core use of _Py_IDENTIFIER() with statically initialized global objects. (gh-30928) We're no longer using _Py_IDENTIFIER() (or _Py_static_string()) in any core CPython code. It is still used in a number of non-builtin stdlib modules. The replacement is: PyUnicodeObject (not pointer) fields under _PyRuntimeState, statically initialized as part of _PyRuntime. A new _Py_GET_GLOBAL_IDENTIFIER() macro facilitates lookup of the fields (along with _Py_GET_GLOBAL_STRING() for non-identifier strings). https://bugs.python.org/issue46541#msg411799 explains the rationale for this change. The core of the change is in: * (new) Include/internal/pycore_global_strings.h - the declarations for the global strings, along with the macros * Include/internal/pycore_runtime_init.h - added the static initializers for the global strings * Include/internal/pycore_global_objects.h - where the struct in pycore_global_strings.h is hooked into _PyRuntimeState * Tools/scripts/generate_global_objects.py - added generation of the global string declarations and static initializers I've also added a --check flag to generate_global_objects.py (along with make check-global-objects) to check for unused global strings. That check is added to the PR CI config. The remainder of this change updates the core code to use _Py_GET_GLOBAL_IDENTIFIER() instead of _Py_IDENTIFIER() and the related _PyId functions (likewise for _Py_GET_GLOBAL_STRING() instead of _Py_static_string()). This includes adding a few functions where there wasn't already an alternative to _PyId(), replacing the _Py_Identifier * parameter with PyObject . The following are not changed (yet): stop using _Py_IDENTIFIER() in the stdlib modules * (maybe) get rid of _Py_IDENTIFIER(), etc. entirely -- this may not be doable as at least one package on PyPI using this (private) API * (maybe) intern the strings during runtime init https://bugs.python.org/issue46541	2022-02-08 13:39:07 -07:00
Pablo Galindo Salgado	69e10976b2	bpo-46521: Fix codeop to use a new partial-input mode of the parser (GH-31010)	2022-02-08 11:54:37 +00:00
Paul m. p. P	89b13042fc	bpo-14916: use specified tokenizer fd for file input (GH-31006) @pablogsal, sorry i failed to rebase to main, so i recreated https://github.com/python/cpython/pull/22190#issuecomment-1024633392 > PyRun_InteractiveOne\() functions allow to explicitily set fd instead of stdin. but stdin was hardcoded in readline call. > This patch does not fix target file for prompt unlike original bpo one : prompt fd is unrelated to tokenizer source which could be read only. It is more of a bugfix regarding the docs : actual documentation say "prompt the user" so one would expect prompt to go on stdout not a file for both PyRun_InteractiveOne\() and PyRun_InteractiveLoop\*(). Automerge-Triggered-By: GH:pablogsal	2022-02-01 14:33:52 -08:00
Pablo Galindo Salgado	a0efc0c196	bpo-46091: Correctly calculate indentation levels for whitespace lines with continuation characters (GH-30130)	2022-01-25 22:12:14 +00:00
Kumar Aditya	41026c3155	bpo-45855: Replaced deprecated `PyImport_ImportModuleNoBlock` with PyImport_ImportModule (GH-30046)	2021-12-12 10:45:20 +02:00
Pablo Galindo Salgado	4325a766f5	bpo-46054: Fix parsing error when parsing non-utf8 characters in source files (GH-30068)	2021-12-12 07:06:50 +00:00
Pablo Galindo Salgado	4f006a789a	Ensure the str member of the tokenizer is always initialised (GH-29681)	2021-11-21 02:06:39 +00:00
Pablo Galindo Salgado	81f4e116ef	bpo-45811: Improve error message when source code contains invisible control characters (GH-29654)	2021-11-20 18:28:28 +00:00
Pablo Galindo Salgado	25835c518a	bpo-45738: Fix computation of error location for invalid continuation (GH-29550) characters in the parser	2021-11-14 01:06:41 +00:00
Pablo Galindo Salgado	cdc7a58277	bpo-45562: Ensure all tokenizer debug messages are printed to stderr (GH-29270)	2021-10-28 18:06:15 +01:00
Pablo Galindo Salgado	10bbd41ba8	bpo-45562: Print tokenizer debug messages to stderr (GH-29250)	2021-10-27 14:27:34 -07:00
Nikita Sobolev	4bc5473a42	bpo-45574: fix warning about `print_escape` being unused (GH-29172) It used to be like this: <img width="1232" alt="Снимок экрана 2021-10-22 в 23 07 40" src="https://user-images.githubusercontent.com/4660275/138516608-fef6ec01-a96a-40f4-81ef-52265b0f536b.png"> Quick `grep` tells that it is just used in one place under `Py_DEBUG`: `f6e8b80d20/Parser/tokenizer.c (L1047-L1051)` <img width="752" alt="Снимок экрана 2021-10-22 в 23 08 09" src="https://user-images.githubusercontent.com/4660275/138516684-ea503136-1e92-48a5-95bb-419e190d5866.png"> I am not sure, but it also looks like a private thing, it should not affect other users. Automerge-Triggered-By: GH:pablogsal	2021-10-22 14:57:24 -07:00
Pablo Galindo Salgado	86dfb55d2e	bpo-45562: Only show debug output from the parser in debug builds (GH-29140)	2021-10-22 01:52:24 -07:00
Victor Stinner	713bb19356	bpo-45434: Mark the PyTokenizer C API as private (GH-28924) Rename PyTokenize functions to mark them as private: * PyTokenizer_FindEncodingFilename() => _PyTokenizer_FindEncodingFilename() * PyTokenizer_FromString() => _PyTokenizer_FromString() * PyTokenizer_FromFile() => _PyTokenizer_FromFile() * PyTokenizer_FromUTF8() => _PyTokenizer_FromUTF8() * PyTokenizer_Free() => _PyTokenizer_Free() * PyTokenizer_Get() => _PyTokenizer_Get() Remove the unused PyTokenizer_FindEncoding() function. import.c: remove unused #include "errcode.h".	2021-10-13 17:22:14 +02:00
Victor Stinner	d943d19172	bpo-45439: Move _PyObject_CallNoArgs() to pycore_call.h (GH-28895) * Move _PyObject_CallNoArgs() to pycore_call.h (internal C API). * _ssl, _sqlite and _testcapi extensions now call the public PyObject_CallNoArgs() function, rather than _PyObject_CallNoArgs(). * _lsprof extension is now built with Py_BUILD_CORE_MODULE macro defined to get access to internal _PyObject_CallNoArgs().	2021-10-12 08:38:19 +02:00
Victor Stinner	ce3489cfdb	bpo-45439: Rename _PyObject_CallNoArg() to _PyObject_CallNoArgs() (GH-28891) Fix typo in the private _PyObject_CallNoArg() function name: rename it to _PyObject_CallNoArgs() to be consistent with the public function PyObject_CallNoArgs().	2021-10-12 00:42:23 +02:00
Noah Kantrowitz	be42c06bb0	Update URLs in comments and metadata to use HTTPS (GH-27458)	2021-07-30 15:54:46 +02:00
Pablo Galindo Salgado	f24777c2b3	bpo-44317: Improve tokenizer errors with more informative locations (GH-26555)	2021-07-10 01:29:29 +01:00
Binbin	17b16e13bb	Fix typos in multiple files (GH-26689) Co-authored-by: Terry Jan Reedy <tjreedy@udel.edu>	2021-06-12 22:47:44 -04:00
Pablo Galindo	a342cc5891	bpo-44396: Update multi-line-start location when reallocating tokenizer buffers (GH-26676) Automerge-Triggered-By: GH:pablogsal	2021-06-12 10:53:49 -07:00
Serhiy Storchaka	2ea6d89028	bpo-43833: Emit warnings for numeric literals followed by keyword (GH-25466) Emit a deprecation warning if the numeric literal is immediately followed by one of keywords: and, else, for, if, in, is, or. Raise a syntax error with more informative message if it is immediately followed by other keyword or identifier. Automerge-Triggered-By: GH:pablogsal	2021-06-08 16:31:10 -07:00
Pablo Galindo	bd7476dae3	bpo-44201: Avoid side effects of "invalid_*" rules in the REPL (GH-26298) When the parser does a second pass to check for errors, these rules can have some small side-effects as they may advance the parser more than the point reached in the first pass. This can cause the tokenizer to ask for extra tokens in interactive mode causing the tokenizer to show the prompt instead of failing instantly. To avoid this, add a new mode to the tokenizer that is activated in the second pass and deactivates asking for new tokens when the interactive line is finished. As the parsing should have reached the last line in the first pass, the second pass should not need to ask for more tokens.	2021-05-22 23:05:00 +01:00
Pablo Galindo	92a02c1f7e	Fix tokenizer error when raw decoding null bytes (GH-25080)	2021-03-30 00:24:49 +01:00
Pablo Galindo	261a452a13	bpo-25643: Refactor the C tokenizer into smaller, logical units (GH-25050)	2021-03-28 23:48:05 +01:00
Pablo Galindo	cd8dcbc851	bpo-43410: Fix crash in the parser when producing syntax errors when reading from stdin (GH-24763)	2021-03-14 04:38:40 +01:00
Batuhan Taskaya	a698d52c39	bpo-40176: Improve error messages for unclosed string literals (GH-19346) Automerge-Triggered-By: GH:isidentical	2021-01-20 13:38:47 -08:00
Pablo Galindo	ae7d3cd980	bpo-42864: Fix compiler warning in the tokenizer with the new paren stack for column numbers (GH-24266)	2021-01-20 12:53:52 +00:00
Pablo Galindo	d6d6371447	bpo-42864: Improve error messages regarding unclosed parentheses (GH-24161)	2021-01-19 23:59:33 +00:00
Lysandros Nikolaou	e5fe509054	bpo-42827: Fix crash on SyntaxError in multiline expressions (GH-24140) When trying to extract the error line for the error message there are two distinct cases: 1. The input comes from a file, which means that we can extract the error line by using `PyErr_ProgramTextObject` and which we already do. 2. The input does not come from a file, at which point we need to get the source code from the tokenizer: * If the tokenizer's current line number is the same with the line of the error, we get the line from `tok->buf` and we're ready. * Else, we can extract the error line from the source code in the following two ways: * If the input comes from a string we have all the input in `tok->str` and we can extract the error line from it. * If the input comes from stdin, i.e. the interactive prompt, we do not have access to the previous line. That's why a new field `tok->stdin_content` is added which holds the whole input for the current (multiline) statement or expression. We can then extract the error line from `tok->stdin_content` like we do in the string case above. Co-authored-by: Pablo Galindo <Pablogsal@gmail.com>	2021-01-14 21:36:30 +00:00
Victor Stinner	00d7abd7ef	bpo-42519: Replace PyMem_MALLOC() with PyMem_Malloc() (GH-23586) No longer use deprecated aliases to functions: * Replace PyMem_MALLOC() with PyMem_Malloc() * Replace PyMem_REALLOC() with PyMem_Realloc() * Replace PyMem_FREE() with PyMem_Free() * Replace PyMem_Del() with PyMem_Free() * Replace PyMem_DEL() with PyMem_Free() Modify also the PyMem_DEL() macro to use directly PyMem_Free().	2020-12-01 09:56:42 +01:00

1 2 3 4 5 ...

316 Commits