pdf.js

Author	SHA1	Message	Date
Jonas Jenwald	3a7fce49a3	A tiny improvement of the `MetadataParser._repair` method We can just insert the initial greater-than sign at the start of the buffer, rather than doing that manually at the end.	2023-02-04 12:43:55 +01:00
Jonas Jenwald	2de03a7d91	Improve how we cache Promises in `WorkerTransport` A number of methods have their Promises cached, to avoid repeated worker round-trips, since they're expected to be called more than once from the default viewer. The way that the caching is currently implemented means that we need to remember to manually clear these Promises on document cleanup/destruction, and it'd be nice to avoid that. With this patch the relevant Promises are now instead placed in just one `Map`, which is easy to clear, and a new helper method is also introduced to reduce duplication for simple `WorkerTransport` methods.	2023-02-04 11:57:37 +01:00
Calixte Denizet	185281957d	[Editor] Make the annotation editor layer invisible when disabled and empty It'll help to avoid to consider them when the browser is restyling.	2023-02-01 17:53:44 +01:00
Jonas Jenwald	cf8ee47589	Remove unused parameters from the `onOpenWithTransport` method in `PDFViewerApplication.initPassiveLoading` The only parameter that we actually need here is the `PDFDataRangeTransport`-instance, since the others are not necessary. - The `url` parameter, as passed to the `getDocument` function in the API, is simply being ignored; see `2d87a2eb1c/src/display/api.js (L447-L458)` - The `length` parameter, as passed to the `getDocument` function in the API, is always being overwritten; see `2d87a2eb1c/src/display/api.js (L519-L525)`	2023-02-01 09:33:22 +01:00
Jonas Jenwald	5e88228767	Allow, optionally, using worker-modules during local development Until PR 12563 is deemed safe to land, I'd still like to be able to use worker-modules in the viewer during local development. Hence this patch which temporarily adds a new `workerModules` hash-parameter, only available in non-PRODUCTION mode, that allows using worker-modules in the development viewer. To enable this functionality, simply use http://localhost:8888/web/viewer.html#workerModules=true	2023-01-31 12:09:44 +01:00
Jonas Jenwald	c5d6391898	[api-minor] Let the `cMapPacked` parameter, in `getDocument`, default to `true` The initial CMap support was added in PR 4259 using the "raw" Adobe files, however they were quickly deemed to be unnecessarily large. As a result PR 4470 introduced the more compact "binary" CMap format, with both of those PRs being included in the very same release (version `0.8.1334`) . Please note that we've thus never shipped anything except the "binary" CMap files with the PDF library, and furthermore note that we've not even once updated the CMap files since they were originally added almost nine years ago. Requiring users to remember that `cMapPacked = true` is necessary, in addition to setting the `cMapUrl` parameter, in order for CMap loading to work feels like a less than ideal API. Hence this patch, which suggests that we simply let `cMapPacked` default to `true` now.	2023-01-30 15:35:02 +01:00
Jonas Jenwald	808ca828f1	Extend `getGlyphMapForStandardFonts` with additional entries (issue 15977)	2023-01-30 12:13:21 +01:00
Tim van der Meij	ee3be2f979	Merge pull request #15951 from Snuffleupagus/polyfill-Path2D Polyfill `Path2D` in Node.js environments	2023-01-28 19:06:54 +01:00
Tim van der Meij	e539d2da1e	Merge pull request #15964 from Snuffleupagus/getDocument-non-object Only accept non-objects passed to `getDocument` in GENERIC builds	2023-01-28 18:42:09 +01:00
Jonas Jenwald	cf0369d622	Polyfill `Path2D` in Node.js environments Until just recently the only existing `Path2D` polyfill didn't have support for Node.js and/or the `node-canvas` package. Given that this was just fixed, in the latest version, we can now finally remove our inline-checks at the relevant call-sites; please also see https://github.com/nilzona/path2d-polyfill#usage-with-node-canvas	2023-01-28 18:28:22 +01:00
Tim van der Meij	cb5a28ceca	Merge pull request #15954 from Snuffleupagus/getDocument-URL-tweaks Tweak the internal handling of the `url`-parameter in `getDocument` (PR 13166 follow-up)	2023-01-28 18:18:17 +01:00
Tim van der Meij	16aef95937	Merge pull request #15966 from Snuffleupagus/GlobalWorkerOptions-defaults Simplify setting the `GlobalWorkerOptions` default values (PR 9480 follow-up)	2023-01-28 18:15:32 +01:00
Calixte Denizet	6f4d037a8e	[JS] Correctly format field with numbers (bug 1811694, bug 1811510) In PR #15757, a value is automatically converted into a number when it's possible but the case of numbers like "000123" has been overlooked and their format must be preserved. When a script is doing something like "foo.value + bar.value" and the values are numbers then "foo.value" must return a number but the displayed value must be what the user entered or what a script set, so this patch is just adding a a field _orginalValue in order to track the value has it has defined. Some people are used to use a comma as decimal separator, hence it must be considered when a value is parsed into a number. This patch is fixing a regression introduced by #15757.	2023-01-26 14:57:02 +01:00
Jonas Jenwald	1c4af2727c	Simplify setting the `GlobalWorkerOptions` default values (PR 9480 follow-up) There's really no need for these "complicated" default value assignments, since `GlobalWorkerOptions` is a local variable at this point, and this is rather a case of too much copy-and-paste. Note that years ago, when all options were set using a global `PDFJS` object, it's possible that options had been set (from the outside) before the object had been properly initialized; see e.g. `a89071bdef/src/display/global.js`	2023-01-26 14:16:01 +01:00
Jonas Jenwald	4758e6649c	Only accept non-objects passed to `getDocument` in GENERIC builds In general it's always recommended to pass a parameter object when calling the `getDocument`-function in the API, since that's the only way to provide additional options, and the fact that it also accepts a URL or TypedArray directly is now mostly for backwards compatibility reasons. Unfortunately we cannot really remove this, since that code has existed since "forever", however we can limit it to only the GENERIC build to avoid completely unnecessary checks in e.g. the Firefox PDF Viewer. Finally, note that the default-viewer always provides a parameter object when calling the `getDocument`-function and it's thus completely unaffected by these changes.	2023-01-26 10:48:58 +01:00
Jonas Jenwald	755319130e	Tweak the internal handling of the `url`-parameter in `getDocument` (PR 13166 follow-up) - Use a `URL`-instance directly, since it's by definition an absolute URL. - Actually limit the "raw" url-string handling to Node.js environments, as intended. - Skip the warning, since we're already throwing an Error if the `url`-parameter is invalid.	2023-01-24 11:18:41 +01:00
Tim van der Meij	edfdb693e5	Merge pull request #15948 from Snuffleupagus/bug-1811668 Tweak `adjustType1ToUnicode` for fonts with a predefined named encoding (bug 1811668, PR 14050 follow-up)	2023-01-21 14:02:40 +01:00
Tim van der Meij	a27d7ba524	Merge pull request #15943 from Snuffleupagus/deprecate-direct-PDFDataRangeTransport [api-minor] Deprecate calling `getDocument` directly with a `PDFDataRangeTransport`-instance	2023-01-21 13:50:20 +01:00
Jonas Jenwald	40a46e4397	Tweak `adjustType1ToUnicode` for fonts with a predefined named encoding (bug 1811668, PR 14050 follow-up) Please note: I cannot reproduce the problem reported in bug 1811668, regarding the context menu, and in any case it's not clear that that part is even a PDF Viewer bug. Looking at bug 1811668 I couldn't help but noticing that the textLayer isn't correct, and it's unfortunately once again a problem with the `adjustType1ToUnicode` function. That's intended to help improve text-selection for fonts without a /ToUnicode-entry, and in many cases it does help (the original PR fixed lots of issues) however it's also caused some problems. In order to improve text-selection in bug 1811668, we'll now properly ignore fonts that have a predefined named encoding specified since that's really the intention with PR 14050.	2023-01-21 12:21:21 +01:00
Jonas Jenwald	f2fce93826	[JBIG2] Ensure that the `decodeInteger` function returns valid integers (issue 15942) The JBIG2 images in this PDF document are corrupt enough that even Adobe Reader warns about it when opening the file. Please note: I don't really know the JBIG2 image format at all, however from a very brief look at the specification it seems that integers should be 32-bit.	2023-01-19 17:14:17 +01:00
Jonas Jenwald	7976fc7851	[api-minor] Deprecate calling `getDocument` directly with a `PDFDataRangeTransport`-instance In general it's recommended to pass a parameter object when calling the `getDocument`-function in the API, since that's the only way to provide additional options, and the fact that it also accepts a URL or TypedArray directly is now mostly for backwards compatibility reasons. However, the `getDocument`-function also accepts a direct `PDFDataRangeTransport`-instance which just seems unnecessary. Please note: The `PDFDataRangeTransport`-implementation was added specifically for the built-in Firefox PDF Viewer, however it's most likely not commonly used by any third-party (given that it requires manual PDF-data loading). Furthermore, the default-viewer always provides a parameter object when calling the `getDocument`-function and it's thus completely unaffected by these changes.	2023-01-19 14:25:55 +01:00
Jonas Jenwald	7b36686fca	Update the year in the `license_header` files	2023-01-18 22:28:18 +01:00
Jonas Jenwald	d6be5141e9	Fallback to using the `name` table to infer the encoding for TrueType fonts missing such data (issue 15910) The relevant TrueType font is missing both /ToUnicode and /Encoding entires, either of which would have prevented the (current) broken textLayer rendering. My first idea was that we could use the `post` table in the TrueType font, see https://developer.apple.com/fonts/TrueType-Reference-Manual/RM06/Chap6post.html, to get the actual glyphNames and amend the fallback ToUnicode-map that way. Unfortunately that didn't work, since the `post` table only contained ".notdef" and "" (i.e. empty string) entries. Instead we try to use the `name` table in the TrueType font, see https://developer.apple.com/fonts/TrueType-Reference-Manual/RM06/Chap6name.html, to determine if the platform is Windows and thus fallback to generate a ToUnicode-map from the `WinAnsiEncoding`.	2023-01-17 16:04:51 +01:00
Jonas Jenwald	d8d5545e03	Merge pull request #15926 from Snuffleupagus/annotation-appearance-stream Ensure that Annotation `appearance`-entries are actually Streams	2023-01-16 15:00:12 +01:00
Jonas Jenwald	cefaecc2e8	Ensure that Annotation `appearance`-entries are actually Streams Note how all over the `src/core/annotation.js`-code we're assuming that if an `appearance`-entry exists it's also a Stream. However, we're not actually checking that thoroughly enough which causes issues in some badly generated PDF documents.	2023-01-16 13:02:53 +01:00
Jonas Jenwald	397f943ca3	[api-minor] Enable transferring of TypedArray PDF data by default (PR 15908 follow-up) This patch removes the recently introduced `transferPdfData` API-option, and simply enables transferring of TypedArray data by default instead of copying it. This will help reduce main-thread memory usage, however it will take ownership of the TypedArrays. Currently this only applies to the following cases: - TypedArrays passed to the `getDocument`-function in the API, in order to open PDF documents from binary data. - TypedArrays passed to a `PDFDataRangeTransport`-instance, used to support custom PDF document fetching/loading (see e.g. the Firefox PDF Viewer). PLEASE NOTE: To avoid being affected by this, please simply copy any TypedArray data before passing it to either of the functions/methods mentioned above. Now that we transfer TypedArray data that we previously only copied, we need to be more careful with input validation. Given how the `{IPDFStreamReader, IPDFStreamRangeReader}.read` methods will always return ArrayBuffer data, which is then transferred to the worker-thread[1], the actual TypedArray data passed to the API thus need to have the same exact size as its underlying ArrayBuffer to prevent issues. Hence we'll check for this and only allow transferring of safe TypedArray data, and fallback to simply copying the data just as before. This obviously shouldn't be an issue in the Firefox PDF Viewer, but for the general PDF.js library we need to be more careful here. --- [1] See `e09ad99973/src/display/api.js (L2492-L2506)` respectively `e09ad99973/src/display/api.js (L2578-L2590)`	2023-01-14 10:39:36 +01:00
Jonas Jenwald	99cfab18c1	Combine the array-like and ArrayBuffer branches, when handling binary data, in `getDocument`	2023-01-13 13:28:44 +01:00
Jonas Jenwald	e09ad99973	Merge pull request #15916 from Snuffleupagus/fetch-transfer [api-minor] Enabling transferring of data fetched with the `PDFFetchStream` implementation	2023-01-13 13:28:12 +01:00
Jonas Jenwald	1362cd91d0	Improve input validation in `PDFDataTransportStream._onReceiveData` (PR 15908 follow-up) The mozilla-central [method `PdfDataListener.readData`](https://searchfox.org/mozilla-central/rev/893a8f062ec6144c84403fbfb0a57234418b89cf/toolkit/components/pdfjs/content/PdfStreamConverter.jsm#207-210) can return `null`, hence it seems like a very good idea to update `PDFDataTransportStream._onReceiveData` to handle that gracefully since the current code will throw in that case. Also, improves the JSDocs for the `PDFDataRangeTransport` class in the API.	2023-01-12 15:24:59 +01:00
Jonas Jenwald	cee97fcd15	[api-minor] Enabling transferring of data fetched with the `PDFFetchStream` implementation Note how in the API we're transferring the PDF data that's fetched over the network[1]: - `f28bf23a31/src/display/api.js (L2467-L2480)` - `f28bf23a31/src/display/api.js (L2553-L2564)` To support that functionality we have the `PDFDataTransportStream`, `PDFFetchStream`, `PDFNetworkStream`, and `PDFNodeStream` implementations. Here these stream-implementations vary slightly in how they handle `ArrayBuffer`s internally, w.r.t. transferring or copying the data: - In `PDFDataTransportStream` we optionally, after PR 15908, allow transferring of the PDF data as provided externally (used e.g. in the Firefox PDF Viewer). - In `PDFFetchStream` we're currenly always copying the PDF data returned by the Fetch API, which seems unnecessary. As discussed in PR 15908, it'd seem very weird if this sort of browser API didn't allow transferring of the returned data. - In `PDFNetworkStream` we're already, since many years, transferring the PDF data returned by the `XMLHttpRequest` functionality. Note how the `getArrayBuffer` helper function simply returns an `ArrayBuffer` response as-is. - In `PDFNodeStream` we're currently copying the PDF data, however this is unfortunately necessary since Node.js returns data as a `Buffer` object[2]. Given that the `PDFNetworkStream` has been, indirectly, supporting transferring of PDF data for years it would seem really strange if this didn't also apply to the `PDFFetchStream`-implementation. Hence this patch simply enables transferring of PDF data, when accessed using the Fetch API, unconditionally to help reduced main-thread memory usage since the `PDFFetchStream`-implementation is used by default in browsers (for the GENERIC build). --- [1] As opposed to PDF data being provided as e.g. a TypedArray when calling `getDocument` in the API. [2] This is a "special" Node.js object, see https://nodejs.org/api/buffer.html#buffer, which doesn't exist in browsers.	2023-01-12 13:59:21 +01:00
Jonas Jenwald	bbe629018d	[api-minor] Add a new `transferPdfData` option to allow transferring more data to the worker-thread (bug 1809164) Also, removes the `initialData`-parameter JSDocs for the `getDocument`-function given that this parameter has been completely unused since PR 8982 (over five years ago). Note that the `initialData`-parameter is, and always was, intended to be provided when initializing a `PDFDataRangeTransport`-instance.	2023-01-10 21:03:44 +01:00
calixteman	fcaeb5db88	Merge pull request #15901 from calixteman/15289_followup Avoid null ExpansionFactor in type1 fonts (follow-up of #15289)	2023-01-07 18:20:31 +01:00
Jonas Jenwald	74e4b515c5	Merge pull request #15897 from Snuffleupagus/issue-15893 Support parsing encrypted documents in `XRef.indexObjects` (issue 15893)	2023-01-07 16:55:41 +01:00
Calixte Denizet	c170245fc0	Avoid null ExpansionFactor in type1 fonts (follow-up of #15289 )	2023-01-07 16:25:24 +01:00
Calixte Denizet	e565e455e2	Set ExpansionFactor to 0.06 when it's equals to 0 in the private dict of CFF fonts	2023-01-07 14:53:13 +01:00
Tim van der Meij	69113f08f2	Merge pull request #15887 from Snuffleupagus/rm-setPDFNetworkStreamFactory Inline the `setPDFNetworkStreamFactory` functionality in `src/display/api.js`	2023-01-07 13:16:23 +01:00
Tim van der Meij	b428824269	Merge pull request #15879 from Snuffleupagus/useWorkerFetch-defaults [api-minor] Improve the `useWorkerFetch` default value checks	2023-01-07 13:13:25 +01:00
Jonas Jenwald	1d5de9f4f4	Inline the `setPDFNetworkStreamFactory` functionality in `src/display/api.js` Given that this is internal functionality, not exposed in the official API, it's not entirely clear (at least to me) why we can't just initialize this directly in `src/display/api.js` instead. When testing both the development viewer and all the ways in which we run tests, everthing still appears to work just fine with this patch.	2023-01-06 13:23:07 +01:00
Jonas Jenwald	7d94fdeb48	Support parsing encrypted documents in `XRef.indexObjects` (issue 15893) Please note: The reduced test-case is not a perfect reproduction of the original PDF document, since this one fails to open in e.g. Adobe Reader, but I do believe that it captures the most important points here. For corrupt and encrypted PDF documents, it's possible that only some trailer dictionaries actually contain an /Encrypt-entry. Previously we'd could easily miss that, since we generally pick the first not obviously corrupt trailer dictionary, and the solution implemented here is to simply pre-parse all trailer dictionaries to see if there's any /Encrypt-entries.	2023-01-06 13:09:37 +01:00
Calixte Denizet	dea2471e96	[JS] UserActivation must be enabled before running document actions else auto-print is broken (it's a regression from patch #15822).	2023-01-04 21:26:36 +01:00
Jonas Jenwald	6bdbb5c5ca	Update the `type`/`subtype` at the end of font parsing This fixes a warning reported by CodeQL, and should also make general sense given that we parse the font-data to determine the actual `type`/`subtype` rather than trusting the PDF document.	2023-01-02 16:21:48 +01:00
Jonas Jenwald	1a69d537c1	[api-minor] Limit the `PDFDocumentLoadingTask.onUnsupportedFeature` functionality to GENERIC builds (PR 15758 follow-up) This was deprecated in PR 15758 but it's unfortunately quite difficult to tell if third-party users are depending on this, e.g. to implement custom error reporting, and if so to what extent. However, thanks to the pre-processor we can limit most of this code to GENERIC builds which still seem like a worthwhile change. These changes reduce the bundle size of the Firefox PDF Viewer by 3.8 kB in total.	2023-01-01 17:53:12 +01:00
Jonas Jenwald	0c1fb4e740	[api-minor] Remove the `PDFDocumentProxy.stats` getter (PR 15758 follow-up) This was deprecated in PR 15758 and given that it's quite unlikely that any third-party users are relying on this functionality, since it was only ever added to support telemetry reporting in the Firefox PDF Viewer, it should hopefully be fine to remove this fairly quickly. These changes reduce the bundle size of the Firefox PDF Viewer by 4.5 kB in total.	2023-01-01 17:06:47 +01:00
Jonas Jenwald	2c57a4232c	[api-minor] Improve the `useWorkerFetch` default value checks Given that the Fetch API only supports the http/https protocols, worker-thread fetching of CMaps and Standard-fonts may thus fail in certain cases. To improve the default behaviour we'll now also check that the `cMapUrl` and `standardFontDataUrl` options are appropriate, except in Firefox where this should always work.	2023-01-01 14:48:28 +01:00
Jonas Jenwald	3110d1f29a	Merge pull request #15869 from Snuffleupagus/_abortOperatorList-clearTimeout Always abort a pending `streamReader` cancel timeout in `PDFPageProxy._abortOperatorList` (PR 15825 follow-up)	2022-12-27 13:26:43 +01:00
Jonas Jenwald	841abb53e6	Remove `PDFPageProxy.getJSActions` caching, since it's unused, in the API Note how, in the scripting initialization in the viewer, we only ever invoke `PDFPageProxy.getJSActions` once per page in order to improve overall performance; see `a575aa13b9/web/pdf_scripting_manager.js (L372-L375)` Hence it really shouldn't be necessary to cache its result in the API, especially when that is done manually rather than using something like `shadow`.	2022-12-27 10:39:33 +01:00
Jonas Jenwald	ae24dbd064	Always abort a pending `streamReader` cancel timeout in `PDFPageProxy._abortOperatorList` (PR 15825 follow-up) When we're destroying a `PDFPageProxy`-instance, during full document destruction, we'll force-abort any worker-thread parsing of operatorLists. Hence we should make sure that any pending cancel timeout is always aborted, since a later `PDFPageProxy._abortOperatorList` call should always "replace" a previous one. Please note: Technically this was always wrong, but with the changes in PR 15825 it became ever so slightly easier to trigger this thanks to the potentially longer timeout.	2022-12-27 10:19:39 +01:00
Jonas Jenwald	2fcf8bb5be	Re-factor searching for incomplete objects in `XRef.indexObjects` (issue 15803) When trying to find incomplete objects, i.e. those missing the "endobj"-string at the end, there's unfortunately a number of possible operators that we need to check for. Otherwise we could miss e.g. the "trailer" at the end of a corrupt PDF document, which is why the referenced document didn't work. Currently we do all searching on the "raw" bytes of the PDF document, for efficiency, however this doesn't really work when we need to check for multiple potential command-strings. To keep the complexity manageable we'll instead use regular expressions here, but we can at least avoid creating lots of substrings thanks to the `RegExp.lastIndex` property; which is well supported across browsers according to https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/RegExp/lastIndex#browser_compatibility Note that this repeated regular expression usage could perhaps be slightly less efficient than the old code, however this method is only invoked for corrupt PDF documents.	2022-12-19 23:01:09 +01:00
Jonas Jenwald	ded02941f2	[api-minor] Move, most of, the `isPureXfa`-handling from `PDFViewer` and into `PDFPageView` By moving this code the "pageviewer"-component example will become slightly more usable on its own, it may simplify a future addition of XFA Foreground document support, and finally also serves as preparation for the following patches.	2022-12-18 13:10:23 +01:00
Calixte Denizet	a84d14b382	[Editor] Avoid to scroll when an annotation is commited (fixes issue #15744 )	2022-12-17 13:48:19 +01:00

... 4 5 6 7 8 ...

5943 Commits